Qwen 3.8 Max API Tested: 16 Thinking Tokens Outperform Disabling Reasoning
Researchers testing Alibaba's Qwen 3.8 Max API found that disabling the model's thinking mode caused accuracy on a two-step math task to drop sharply from 4 out of 4 to just 1 out of 6. Setting a minimal thinking budget of only 16 tokens restored full accuracy at a lower cost than the default configuration. The model's seven reasoning effort levels were found to collapse into just four distinct behaviors based on token caps, with low-effort settings causing incorrect answers on complex tasks. Testing also revealed that the implicit cache requires a prompt of roughly 4,300 tokens before activating, and that context window limits are enforced strictly with clear error responses. The evaluation was conducted on launch day via the Synthorai gateway, as Alibaba released no public model card or benchmark data alongside the model.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in