Researchers use 1970s-style MUD game to benchmark LLMs for just $99
A small team of independent researchers spent several months testing whether a MUD — a text-based game format dating to the 1970s — could serve as a benchmark for evaluating large language models, spending only $99 in API credits. The experiment scored multiple LLMs across four behavioral dimensions, two of which relied on an LLM-based classifier as a judge. When those two classifier-dependent dimensions were removed, one leading model dropped six places in the rankings, raising questions about judge reliability. Cross-checking the classifier against a second judge revealed per-model agreement ranging widely from 85% to just 22%, with an aggregate kappa of 0.04 suggesting significant noise in the instrument. The team acknowledges the work is a proof of concept with notable limitations, and has made all data, code, and transcripts publicly available as they plan a more rigorous Phase 2.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in