Meituan updates LongCat Avatar to v1.5 with faster 8-step inference and INT8 support
Meituan's LongCat team released version 1.5 of its open-source talking-avatar model on May 21, 2026, introducing several technical improvements over the December 2025 original. The update replaces the Wav2Vec2 audio encoder with Whisper-Large-v3, which the team says produces smoother lip synchronization. A new distillation flag reduces generation to just 8 sampling steps, while an INT8 quantization option lowers VRAM requirements, making the 13.6-billion-parameter model more accessible on consumer GPUs. The release also introduces Cross-Chunk Latent Stitching, a technique that eliminates redundant encoding between video segments to enable seamless long-form output. Meituan claims the model outperforms rivals including HeyGen and Kling Avatar 2.0 in internal human evaluations, though no independent benchmark has yet verified these results.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in