7.5B Gemma model outscores 24B Devstral in independent 56-task coding benchmark
A developer ran a custom 56-task coding benchmark across 16 local AI model configurations on a single 16 GB GPU card, completing 36 full runs under identical conditions. Google's Gemma-4-e4b, a 7.5B-parameter model in a 4.97 GiB file, scored 42 out of 56, outperforming Devstral-small-2-24b, a 24B model, which scored 40. All tasks were graded using hidden deterministic tests run by compilers and test runners, with no model judging another model's output. The results did not align with any public leaderboard and in at least one case directly inverted a public ranking. The benchmark also revealed significant score variance between identical consecutive runs for some models, suggesting output consistency is itself a meaningful performance dimension.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in