Benchmark Tests Jev-1.13.0 Decision Model Across 10 Datasets for AI Agent Use
A black-box engineering evaluation of TypeSafe AI's Jev-1.13.0 decision model was conducted in September 2026, spanning 10 public datasets, roughly 22,500 API calls, and 52.2 million input tokens at a total cost of $2.19. Unlike large language models, Jev accepts a state and typed questions, returning typed answers with probabilities and confidence scores without generating free-form text. The model performed strongly on tasks such as indirect prompt-injection detection, document reranking, intent classification, and tool routing, making it a viable fast-decision layer in agent systems. However, it showed no meaningful signal for predicting another model's failure modes, trajectory failure attribution, and performed notably worse on non-English tasks due to its English-primary training. Two key engineering findings emerged: decomposing fuzzy judgments into orthogonal sub-questions and using a two-stage choice-then-verify approach each dramatically improved accuracy in their respective tasks.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in