Claude Sonnet 5.5 vs Opus 5.5 vs GPT-6.1 Sol

Claude Sonnet 5.5 edged Claude Opus 5.5 on our general-programming benchmark, a payments app, 0.7984 to 0.7926, a margin I read as level. It built the better app, though: it earned 0.9894 to Opus's 0.8732 before a ceiling squeezed them together. Then Opus 5.5 outclassed it on an Atlassian Forge app, 0.9767 to 0.5508. Same two models, same goose, same 150-call budget, same week. The order flipped and the gap went from 0.0058 to 0.4259.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in