Glasshouse v0.1 Launches as Open Memory Benchmark for AI Systems
Glasshouse v0.1 is a newly released open-source long-term memory benchmark designed to test AI systems more rigorously than existing vendor-published metrics. The benchmark comprises 2,847 questions spanning a 1.97-million-token conversation in 10 languages, with 50 photographs and conversation sizes ranging from 1,882 to 103,572 turns. Unlike single-score benchmarks, Glasshouse reports results across separate axes — including recall, stale fact handling, contradiction detection, and false memory — since a system can perform well on one while failing another. It also rewards appropriate uncertainty, scoring an 'I don't know' response higher than a confidently wrong answer when facts have changed or conflict. The project is publicly available on GitHub and open to submissions from both individuals and companies, with community feedback already having shaped several of its core evaluation axes.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in