AI Memory Benchmark Site Bench'd Accused of Flawed Scores and Broken Verification
A developer auditing Bench'd, a paid AI memory benchmarking platform, found that the top-ranked scores on its leaderboard do not match the underlying data in its own repository, as reviewed on August 29, 2026. The site's top three leaderboard entries show mathematical inconsistencies, including a perfect score paired with a low reliability rating and a non-zero score derived from all-zero dimensions. Bench'd's published self-verification method relies on a domain, benchd.dev, that does not exist, meaning no receipt has ever been independently verifiable using the site's own instructions. The harness repository, which underpins the platform's claims of open and reproducible benchmarking, has had no code commits in over 80 days and has five unresolved issues from vendors. These findings were filed as a GitHub issue on August 23, 2026, and the contested scores remain live on the leaderboard while the platform continues to charge vendors up to $3,999.99 per month for a verification badge.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in