LeaDQ Uses Reinforcement Learning to Select Unlabeled Samples for Federated Training
A research paper introduces LeaDQ, a method designed to solve the data-querying problem in federated learning, where clients continuously receive unlabeled data streams. Because labeling every incoming sample is costly, clients must decide which samples are worth spending their limited annotation budget on. The core challenge is that samples most useful to a local client may not be the most valuable for optimizing the shared global model. LeaDQ frames this selection process as a reinforcement learning problem, where a policy makes binary decisions on each incoming sample using the model's predictive logits as observations. Selected samples are sent to an expert for labeling, incorporated into federated training, and the cycle repeats as new unlabeled data arrives.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in