Qwen3.8 Flash Outperforms More Expensive Max Prime in Atlassian Forge Benchmark

In benchmark tests, the smaller Qwen3.8 Flash model unexpectedly outperformed the larger Max Prime model on an Atlassian Forge application. The cheaper Flash model scored 0.7961 while Max Prime scored 0.6953 on the Forge benchmark, though Max Prime later outperformed Flash on a separate payments application test. Max Prime's lower Forge score resulted from a single rule violation involving storage index attributes, which incurred a scoring penalty. Both models demonstrated they could follow the correct rule when explicitly prompted but did not inherently know it. The tests were conducted using the same agent framework with a 150-call budget and no internet access.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in