ServiceNow Releases EVA-Bench Data 2.0 with 213 Voice Agent Test Scenarios Across 3 Domains
ServiceNow AI Research has launched EVA-Bench Data 2.0, an open-source benchmark designed to evaluate enterprise voice agents across airline customer service, IT service management, and healthcare HR service delivery. The updated benchmark covers 213 evaluation scenarios and 121 tools — approximately four times the scope of its predecessor. All scenarios were derived from real phone-based customer service workflows and validated for solvability by three leading AI models: OpenAI GPT-4.5, Google Gemini 3.1 Pro, and Anthropic Claude Opus 4.6. The healthcare domain notably incorporates domain-specific regulatory details such as NPI provider identifiers, FMLA regulations, and insurance coverage rules to reflect real-world complexity. The datasets are freely available on Hugging Face, with a multilingual expansion planned for a future release.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in