Gemma 4 26B Trails Jev by 2.1 Points as a Decision Model on Single AWS L4 GPU
A pre-registered benchmark tested Google's open-source Gemma 4 26B model as a decision classifier on a single AWS EC2 L4 GPU, comparing it against TypeSafe's hosted Jev 1.13.0 service and Google's DiffusionGemma variant. Evaluated on Bespoke Labs' 3,880-record public test suite, plain Gemma 4 26B scored 2.1 percentage points below Jev overall, with no meaningful gap on yes/no questions but a 4.5-point deficit on multiple-choice items. Jev showed better out-of-the-box calibration, though fitting a single temperature parameter on just 50 labels brought Gemma's median calibration error to within 0.01 of Jev's. Compared to DiffusionGemma, the plain Gemma read matched accuracy, performed similarly after calibration fitting, and ran 1.9 to 5.1 times faster per decision. All per-item outputs and pre-registration documents were committed to a public GitHub repository before any model calls were made.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in