Study Tests 8 LLMs on Informal Indonesian Customer Chat for Small Businesses

A developer who builds WhatsApp bots for Indonesian small businesses (UMKM) created a benchmark to test how well large language models handle the way customers actually type online. The benchmark, called UMKM-Bench, used three fictional shops set in Yogyakarta, Bandung, and Semarang, each with its own product catalog, shipping table, and policies. A total of 84 customer messages were written across three versions — formal, slang, and slang with typos — drawing on Javanese and Sundanese informal language mixed with abbreviations and English loanwords. Eight LLMs were evaluated on two key abilities: understanding garbled real-world queries and avoiding hallucinated answers when the shop's data could not support a response. Scoring was handled with plain Python, checking whether models correctly identified intent, product, quantity, and city, while penalizing any fabricated figures not found in the shop data or the customer's message.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in