Outdated Agent Prompts, Not Newer AI Models, May Be Behind Performance Drops
A developer noticed that coding agent workflows felt worse after upgrading to newer AI models like Claude Opus 5 and GPT-5.6, initially blaming the models themselves. Prompted by Andrej Karpathy's insights on agent failures, the developer audited existing custom skills and found many instructions were written around older model behaviors, causing redundancy or over-execution. Cleaning up stale and conflicting prompts led to noticeably improved agent performance. Both Anthropic and OpenAI officially recommend recalibrating instructions during model migrations, with OpenAI noting leaner system prompts improved internal coding-agent scores by roughly 10–15%. An ICLR 2024 study also found that prompt performance correlates only weakly across different models, reinforcing that prompts require regular maintenance as underlying models evolve.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in