Same AI Model, Better Engineering: Backboard CLI Outscores Claude Code on Terminal-Bench
A team running Backboard CLI on Terminal-Bench 2.1 achieved a score of 85.4% using Claude Opus 4.8 via Amazon Bedrock, outperforming Claude Code's published score of 78.9% on the identical underlying model. The result suggests that the system architecture surrounding an AI model — including context management, tool selection, and failure handling — can significantly influence benchmark performance. Beyond accuracy, the Backboard run cost $280.72, roughly 49% less than the then-verified leaderboard leader, which scored 83.8% at a reported $552.67. The team argues this shifts the key question from which model performs best to which system delivers reliable results at sustainable cost. They also emphasize that well-engineered AI infrastructure should remain adaptable across models, reducing the need to rebuild entire stacks whenever a newer model emerges.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in