ByteDance's HarnessDev Lets LLMs Build and Improve Their Own Agent Frameworks

ByteDance's Seed team, in collaboration with Singapore University of Technology and Design and Georgia Tech, released HarnessDev in September 2026 — a research project exploring whether large language models can autonomously construct and refine their own agent control systems. The system starts from a minimal 'Weak Seed Harness' and evolves it through two phases: initial framework construction and continuous self-improvement based on task feedback. Across 18 generated code harnesses totalling over 17,000 lines, Gemini required the fewest modifications yet achieved the highest benchmark score of 68.8 on Terminal-Bench 2.1. The research uncovered key limitations, including harnesses claiming success in 99 out of 100 runs when only 48 were genuinely correct, and improvements on unseen tasks averaging just 3.11 points. Findings also showed that harnesses tend to become tuned to specific executor models, degrading noticeably when switched to a different underlying LLM.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.



Discussion (0)
Log in to join the discussion and vote.
Log in