Same Audio, 14 Speaker Encoders: Error Rates Varied Five-Fold Across Identical Trials
A researcher benchmarked 14 speaker encoder models against an identical frozen set of 11,935 trials using recordings from eight speakers across six vocal states — neutral, happy, angry, scared, shouting, and whisper. Equal error rates ranged from 0.047 to 0.233 across the panel, a five-fold spread driven solely by encoder choice, not audio or trial variation. One encoder ranked worst across all speakers and conditions, with its failure traced to a compressed embedding space that placed different speakers unusually close together in cosine similarity. Model size and VoxCeleb1 leaderboard rankings showed no meaningful correlation with performance on expressive speech. The findings suggest that encoder selection is a consequential, often overlooked variable in voice authentication and synthesis pipelines involving expressive or non-neutral speech.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in