How strict audio validation silently killed emotion in a TTS training pipeline
A developer building a training dataset for an emotion-expressive text-to-speech model discovered that their automated quality control process was systematically filtering out the most emotionally rich audio clips. The pipeline used OpenAI's Whisper to validate generated speech against scripts, but emotion-laden audio — featuring trembling, pitch variation, or fading — was harder for Whisper to transcribe accurately and failed validation more often. As a result, flat, monotone takes consistently passed the filter, biasing the training corpus toward emotionless speech. Cosine similarity measurements confirmed the problem: bulk-generated clips scored 0.77–0.94 against a neutral voice, while a well-crafted emotional sample scored just 0.164. The fix involved a two-stage selection process — first filtering for transcription accuracy, then ranking surviving clips by emotional distance from neutral to pick the most expressive take.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in