Active Learning Cuts Labelling Costs but Only Works Under Specific Conditions
Active learning is a machine learning technique where a model selects which unlabelled examples should be annotated next, prioritising the most informative ones rather than sampling randomly. Core strategies include uncertainty sampling, query by committee, diversity-based core-set selection, and expected model change estimation, with the most practical approach combining uncertainty and diversity. A key failure mode occurs when uncertainty-ranked queues surface genuinely ambiguous, mislabelled, or invalid data, sending annotators items that cannot be correctly labelled by anyone. Filtering the pool for basic validity before ranking is recommended to address this problem. Literature suggests active learning can reduce labelling requirements by around 40% in optimal conditions, though this figure is considered optimistic and depends heavily on the data pool containing a meaningful proportion of hard, informative examples.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in