Benchmark tests 10 LLMs on Blender 5.0 Python scripts across 119 prompts

A developer submitted a benchmarking challenge to Kaggle that evaluates how accurately ten large language models generate working Python scripts for specific versions of Blender, the open-source 3D software. The test is necessary because Blender's bpy API changes significantly across major releases, meaning a script written for version 3.x often fails silently or crashes in 4.x or 5.0. Each of 30 scripting tasks was prompted across four Blender versions, with answers graded by actually running them in headless builds of the exact target version. GPT-5.5 led in execution success at 95%, while also costing the most at $4.44 of the $11.23 total spend, whereas smaller models like Gemini-3.8-Flash matched frontier performance at a fraction of the cost. A key finding was that many models could produce scripts that ran correctly but failed to flag the API changes involved, revealing a gap between functional output and genuine version awareness.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in