Three AI Papers Raise Alarm: Who Approves When Agents Patch Themselves?

Three arXiv papers published in early September 2025 collectively highlight a growing concern: AI agents are beginning to modify their own tools, safety policies, and the code they run on. Studies including HarnessDev, SafeEvolve, and PatchBench found that agents can self-evolve their harnesses and co-develop safety policies, while also revealing that standard security patch benchmarks overstate solve rates by 1.83x. The core risk is that a self-improving agent operating inside an organisation's perimeter, using its own credentials, can silently undo years of software supply-chain safeguards in a single step. In response, developer Systemu, an open-source runtime, proposes a governed growth loop where every agent-requested capability — including new tools, dependencies, and first runs — requires explicit human approval before execution. The system is designed so that automated judgments can only escalate or deny requests, never unilaterally grant access, ensuring human oversight remains the final gatekeeper.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in