Study finds AI models misuse German legal and technical terms in professional contexts
A developer testing AI language models found that they frequently use incorrect or outdated German legal and technical terminology, sometimes with serious implications — such as translating 'execution' as 'Hinrichtung' (putting to death) in a project plan. To evaluate this systematically, the developer built a terminology benchmark using 11 vehicle-inspection tasks with 47 expected terms, each traceable to a cited provision of German or EU law. Results showed GPT-6 Astra led with 83% term coverage, while Claude Opus 5 scored lowest at 64%, though five of the nine models tested fell within three terms of each other. The developer also noted that an earlier version of the benchmark was inadvertently biased because the expected terms were drafted with help from the same vendor whose models ranked highest. The findings highlight that fluency in a language does not guarantee domain-specific or jurisdiction-specific accuracy, and that benchmark design itself can significantly skew comparative results.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in