Why AI Models Process Hindi and Tamil Far Less Efficiently Than English
AI language models tokenize text differently across languages, and Indian languages like Tamil and Hindi require significantly more tokens than English to express the same sentence. A test comparing five languages found that while English needed about one token per word, Tamil required nearly eleven, making Indian-language AI interactions costlier and slower. This disparity stems from how tokenizers are trained using an algorithm called Byte Pair Encoding, which learns efficient text chunks based on frequency in training data. Since most AI training datasets are heavily English-dominated, Indic scripts receive far fewer opportunities to form compact, efficient tokens. The result is that Indian-language users face higher processing costs and slower responses for conveying the same information an English speaker can express more cheaply.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in