BEACON Method Cuts RL Training Collapse on Long-Horizon AI Agents
Researchers from ZJU-REAL have introduced BEACON, a milestone-guided policy learning framework designed to address a well-documented failure in reinforcement learning for long-horizon language agents. Standard trajectory-level RL methods like GRPO struggle on extended tasks because correct actions receive conflicting gradient signals depending on overall trajectory outcome, and partial successes are treated the same as total failures, wasting over 73% of training samples. BEACON addresses this by splitting trajectories at verifiable milestone states — observable environment transitions requiring no human annotation — and computing rewards and advantage estimates within each segment separately. On the ALFWorld benchmark, the method improved long-task success rates from 53.5% to 92.9% and effective sample utilization from 23.7% to 82.0%. Notably, performance gains increase with task horizon length, suggesting the approach becomes more valuable as task complexity grows.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in