Gemma 4 26B Trails Jev by 2.1 Points Overall in Benchmarked Decision Model Test
A pre-registered benchmark tested Google's open Gemma 4 26B model as a decision model on a single AWS EC2 L4 GPU, comparing it against TypeSafe's hosted Jev 1.13.0 service and Google's DiffusionGemma variant. On Bespoke Labs' 3,880-record public evaluation suite, plain Gemma 4 26B scored 2.1 points below Jev overall, performed comparably on yes/no questions, but lagged 4.5 points behind on multiple-choice tasks. Jev showed better out-of-the-box calibration, though a single temperature adjustment brought Gemma's median calibration error to within 0.01 of Jev's. Against DiffusionGemma, the plain Gemma read matched accuracy levels while running 1.9 to 5.1 times faster per decision. All per-item outputs and pre-registration documents were committed to a public GitHub repository before any model calls were made.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in