Open-source tool distills smaller AI models to cut agent inference costs by 40%
Experiential Labs has launched 'wmo serve', an open-source tool designed to reduce the cost of running AI agents by routing repetitive tasks to smaller, distilled models instead of expensive frontier models. The tool ingests existing agent traces and uses them to continuously train specialized smaller models through distillation from open-source alternatives. A built-in router dynamically decides which tasks require a frontier model and which can be handled by the cheaper distilled model, while token compaction further reduces costs by removing noise. Users can run it locally via an OpenAI-compatible endpoint using an OpenRouter key, or opt for a hosted solution that promises equivalent quality at over 40% lower cost. The project is open-source on GitHub, and a hosted waitlist is available at experientiallabs.ai.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in