Developer Benchmarks 2019 GPT-2 XL Against 2025 Qwen3 on Local AMD Hardware

A developer with no prior experience running local AI models set out to compare OpenAI's 2019 GPT-2 XL against the 2025 Qwen3-0.6B, running both on a personal laptop using AMD's Lemonade local AI server. The experiment, dubbed 'Second Squeeze,' aimed to measure how much AI capability had changed in six years despite the newer model being less than half the size of the older one. An initial speed measurement showed a 297x performance gap, but a reproducibility check revealed the figure was flawed due to cold-start conditions, with the corrected gap settling at approximately 11x. The developer also encountered a dead end attempting to replicate GPT-2's original 2019 accuracy benchmark, as the local server did not return the probability data the standard evaluation tool required, prompting an open-source issue filing. The project highlighted lessons in reproducibility and iterative problem-solving over raw benchmark results.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in