SShortSingh.
Back to feed

Claude Sonnet 5.5 vs Opus 5.5 vs GPT-6.1 Sol

0
·7 views

Claude Sonnet 5.5 edged Claude Opus 5.5 on our general-programming benchmark, a payments app, 0.7984 to 0.7926, a margin I read as level. It built the better app, though: it earned 0.9894 to Opus's 0.8732 before a ceiling squeezed them together. Then Opus 5.5 outclassed it on an Atlassian Forge app, 0.9767 to 0.5508. Same two models, same goose, same 150-call budget, same week. The order flipped and the gap went from 0.0058 to 0.4259.

Read the full story at DEV Community

This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)

Log in to join the discussion and vote.

Log in

Related stories

0
ProgrammingDEV Community ·

"The Law": Enforcing Deterministic Boundaries on AI Tools

In Part 3 of this series, we looked at how to control budget for Agents. raw, unchecked tool execution. Giving an LLM the ability to invoke tools is essentially handing an untrusted natural language parser the keys to your internal infrastructure: databases, email APIs, cloud storage, and payment gateways. If your tool executor blindly trusts what the model outputs, you don't just have a hallucination problem — you have a deterministic security vulnerability. Enter the third member: The Law.

0
ProgrammingDEV Community ·

FareTrace

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend Travelers can use FareTrace, a flight-fare tracking application, to decide whether to book a flight now or wait for a better deal. The idea came from a simple problem. One of our friends travels frequently and often finds it difficult to decide whether the fare they are seeing is actually a good deal. A flight might look cheap at first, but without knowing its previous prices, baggage allowance, number of stops, or duration, it can be difficult to judge the fare properly. We wanted FareTrace to address that proble

0
ProgrammingDEV Community ·

Hacktoberfest Weekend Challenge: Build for a Friend](https://dev.to/challenges/hacktoberfest-weekend-2026-10-01)

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend I built Meydan (ميدان), a bilingual Arabic and English practice studio for people who know their field but struggle to express their knowledge naturally in Arabic. I built it for a friend who studies Arabic and wants to use it beyond ordinary conversation: to discuss professional topics, travel experiences, technical work, and real-world ideas. The problem is familiar to many Arabic learners in non-Arabic-speaking environments: they may study grammar, morphology, literature, and conversation, yet still feel unable

0
ProgrammingDEV Community ·

My sister drives to her patients. Her notes about them never leave her laptop.

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend. My sister Andrea is a speech-language therapist in Colombia who treats patients at home: children audiorapy helps with that: Families book with buttons. "Hola" → consent → open slots → tap one → booked. Reminders go out 24 h before with Confirm / Reschedule / Cancel. Silence is flagged, never acted on.