GPT-6 Astra dominates easy robot task but ties Claude on harder one, tests show
Independent evaluator Robocurve tested OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 on physical robot arms in September 2026, publishing results on consecutive days. Astra completed a simple block-picking task 19 out of 20 times, while Fable 5.1 managed only 8 out of 20. However, both models performed identically on the harder puzzle-placement task, each succeeding just 2 out of 20 trials. Astra also proved faster and cheaper per run on the block task, averaging 2.5 minutes at $0.94 compared to Fable 5.1's 6.8 minutes at $2.12. Robocurve, a Y Combinator-backed Public Benefit Corporation, used an open-source evaluation framework with standardised hardware and human graders, allowing results to be independently reproduced and verified.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in