Six LLMs Tested on AI Workflow Recovery Decisions; Four Score Perfect 100%
A developer built the Paused Workflow Recovery Benchmark to evaluate how large language models handle paused automated workflows across 30 synthetic scenarios with six possible recovery actions. The benchmark was run on September 26, 2026, testing six models from Google, Anthropic, and OpenAI. Four models — Claude Opus 4.8, Gemini 3.1 Pro Preview, GPT-6 Astra, and Gemini 3.7 Flash — achieved perfect 100% accuracy, while Claude Haiku 4.5 and Claude Sonnet 4.5 each scored 93.33%, both misclassifying the same two scenarios. The two lower-scoring models incorrectly chose RESUME over RETRY in cases where the scenario did not explicitly state whether a temporary fault had been resolved, suggesting sensitivity to ambiguous phrasing. Run costs across all models were minimal, ranging from roughly $0.01 to $0.13 per full evaluation.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in