Smaller Typed-Decision Model Outperforms 35B LLM on 12,000 Federal IT RFQ Classifications

A benchmarking study tested three AI models — Jev, Qwen3.5-35B-A3B, and Laya 421M — on 12,000 real U.S. federal IT procurement solicitations to assess how safely each could automate classification tasks for a federal IT reseller. Jev, a typed-decision API, achieved the highest primary-class accuracy at 91.9%, narrowly beating the much larger 35B LLM Qwen at 89.6%, while the 421M open-weight Laya model trailed at 78.0%. The decisive factor was not raw accuracy but calibration — Jev's low calibration error of 0.049 allowed it to reliably flag uncertain predictions, enabling 86.5% of cases to be auto-accepted at a 96.7% precision rate. Qwen failed to reach the confidence thresholds needed for bounded automation, and Laya's high calibration error of 0.322 made safe automation targets unreachable. The findings suggest that for operational automation, a model's ability to signal its own uncertainty can matter more than its headline accuracy score.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in