Laya offers a fast, local 322M-parameter alternative to LLM-based text classification

Laya is an open-source, non-autoregressive decision engine designed to replace costly large language model calls for routine classification tasks such as routing, moderation, and escalation. Unlike standard LLM-as-a-judge setups, Laya runs a single encoder forward pass to answer typed questions — including label selection, scoring, and yes/no probabilities — without generating any tokens. The project ships three model checkpoints on Hugging Face, including a 322M-parameter multilingual model covering over 100 languages, and reports inference speeds as fast as 7.2ms per question when batched on a T4 GPU. Laya runs entirely locally with no API keys or GPU required, needing roughly 1.5GB of disk space and 1GB of RAM. Released on September 18, 2025, the project accumulated over 26,000 GitHub stars within nine days, reflecting widespread frustration with the latency and cost of using frontier models for simple classification decisions.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in