Kaggle AI models tested on bilingual JSON patch contracts, show varying accuracy
A developer conducted a benchmarking challenge on Kaggle on October 1, 2026, to test AI models' ability to produce accurate JSON responses. The test used 'Bilingual Patch Contracts' to evaluate failures where JSON parses correctly but contains wrong data or is wrapped in disruptive Markdown. Three lightweight models from different providers were tested against 36 prompts involving state updates in English, Chinese, and mixed-language instructions. Results showed perfect accuracy for Gemini 3.7 Flash, 66.7% for GPT-5.4 nano, and complete failure for Claude Haiku 4.5, which wrapped all outputs in Markdown.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in