LongHorizon-Harness Lets AI Agents Run Complex Tasks for Hours Without Losing Progress
LongHorizon-Harness is an open-source framework designed to solve a core limitation of AI agents: their inability to sustain long, complex tasks without losing context or failing mid-way. The tool wraps existing agents in a persistent execution loop built around four mechanisms — fresh-context execution, durable verified state, checkpoint-based recovery, and independent auditing. Rather than replacing or retraining models, it gives agent backends like Claude Code or DeepSeek a stable shell that resumes from the last verified step after a failure, instead of restarting from scratch. The project, currently at v0.1.x with an MIT licence and around 1,500 GitHub stars, is benchmarked against WeaveBench, OSWorld 2.0, and Terminal-Bench 2.1, with results detailed in an accompanying arXiv paper. It supports four model backends and is aimed at developers running long-horizon tasks, not casual end users.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in