OpenAI Playbook Calls for Full Harness Disclosure in AI Model Evaluations
OpenAI has released an official playbook urging that third-party evaluations of frontier AI models disclose the full testing environment, not just benchmark scores. The company argues that factors such as prompts, tool access, state management, compute budgets, and scoring mechanisms can materially alter observed model performance. This concern is illustrated through GPT-5.5 cyber-range tasks, where harness features like state compaction visibly shifted results. OpenAI identifies additional risks to evaluation validity, including reward hacking, data contamination, and sandbagging, and recommends these be reported alongside results. The guidance is framed as a transparency standard rather than a universal configuration, aiming to help readers judge whether evaluation conditions actually match the claims being made.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in