How Microsoft's LLMLingua Shrinks AI Prompts Up to 20x by Deleting Predictable Tokens
Microsoft Research developed LLMLingua, a prompt compression tool that uses a small language model to identify and delete tokens a larger model could predict on its own, based on a measure called perplexity. The system applies different compression ratios to different prompt sections, preserving dense instructions while aggressively trimming repetitive few-shot demonstrations. Despite producing grammatically broken, telegraphic text, compressed prompts perform surprisingly well with large language models, which reconstruct missing context without explicit instruction. Benchmarks like GSM8K and BBH showed up to 20x compression with minimal drop in task performance, particularly for reasoning and in-context learning tasks. A second-generation version replaced the causal language model with a bidirectional encoder trained on GPT-4-compressed data, making the process faster and better at judging token importance using both preceding and following context.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in