How Long-Running AI Agents Work for Days: Model Training and Task Orchestration Explained

AI models such as Kimi K3 and Google's Gemini 3.8 Flash are now specifically designed and trained to handle tasks spanning hours or even days, rather than single chat sessions. Kimi K3, a 2.8-trillion-parameter open-weight model by Moonshot, uses reinforcement learning with trajectories of up to one million tokens per round and a checkpoint system that pauses and resumes sandboxes in milliseconds. Research from METR shows that leading models' "time horizon" — the task length completed successfully half the time — has doubled every seven months since 2019, with top models now handling roughly 17-hour tasks. Google's Gemini 3.8 Flash, released on September 2, 2026, is positioned as a workhorse for long-horizon software engineering and agentic tasks, scoring 89.4% on Terminal-Bench 2.1 at a fraction of the cost of rivals. Experts note that model capability alone is only half the equation — the "harness" infrastructure that manages context, checkpointing, and task handoffs is equally critical to making long-running agents reliable.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in