Developer Builds Quiz Pitting Humans Against AI on Expert-Level Academic Questions

A developer has created 'Humans vs. HLE', a web-based quiz that challenges users to compete against frontier AI models on Humanity's Last Exam, a benchmark of expert-level academic questions designed to stump even advanced AI systems. Players answer multiple-choice questions drawn from the HLE dataset, and are eliminated after four wrong answers. Scores are displayed alongside published AI model accuracy rates, and a persistent leaderboard tracks how human participants perform over time. The app runs on Cloudflare Workers and uses AES-GCM encryption and HMAC-signed result tokens to keep answer keys server-side, preventing users from cheating via browser tools. The project was built using Claude and tracked with a tool called Entire, with 140 curated questions bundled directly into the application.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in