Meta's 7B-Parameter ASR Model Supports 1,693 Languages With Under 10% Error Rate
Meta has developed a 7.8 billion-parameter automatic speech recognition model called Meta-Omnilingual-ASR-7B, maintained on Replicate by Subformer, capable of transcribing audio in 1,693 languages. The model combines wav2vec2 feature extraction with an LLM-based decoder, enabling zero-shot and few-shot multilingual transcription without language-specific training data. It achieves character error rates below 10% for 78% of supported languages and delivers near real-time performance on 30-second audio clips, though it requires around 17GB of VRAM. Key use cases include endangered language preservation, multilingual media subtitling, academic speech research, and heritage audio digitization. The model caps standard audio input at 40 seconds, with a separate unlimited-length variant available for longer recordings.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in