OpenAI Token Costs Vary Sharply by Language: Czech 2x, Chinese Near English Parity

A solo developer in Korea tested the same customer-support prompt across 27 languages using OpenAI's tokenizers and found significant cost differences between languages. Simplified Chinese tokenized at nearly the same cost as English (1.03x), while Czech proved the most expensive at 2.0x English token count. The newer o200k_base tokenizer, used since GPT-4o, dramatically reduced costs for Indic, Arabic, and Thai scripts compared to the older cl100k_base encoder — Bengali dropped from 6.1x to 1.68x English tokens. Heavily inflected Latin and Cyrillic-script languages like Polish and Ukrainian now rank among the costlier options due to diacritics splitting into multiple tokens. The developer recommends writing system prompts in English and reserving non-English text only for user-facing content to reduce compounding token costs at scale.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in