JevBench launches open benchmark for typed decision models ranking speed, cost, accuracy
A developer has released JevBench, an open-source benchmark designed to evaluate typed decision models — systems that return bounded choices and probabilities rather than free-form text. Unlike large language models, these so-called Jev-class models are claimed to be significantly faster and cheaper while maintaining comparable intelligence on text inputs. JevBench scores models across four weighted dimensions: chance-corrected intelligence, calibration, speed, and cost, using a standardized set of 534 English-language decision prompts. The current leaderboard is topped by Jev at 74.4, followed closely by SemIf, djev, Winnow-12B Q8, and reflex 4B. The project is available on GitHub, though it carries noted limitations including English-only support and potential noise in scores separated by roughly one point.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in