Filtering Out Unstable Audio Clips Fixes Hoarse TTS Style Synthesis
A text-to-speech system that builds emotional style vectors from a small set of representative audio clips was producing consistently hoarse output due to low-quality clips being included in the selection. Because only five clips are used per style, even one rough clip contributes 20% weight to the average style vector, pulling it toward abnormal acoustic regions. The developer identified two key audio quality factors — jitter (unstable vocal cord vibration) and octave jumps (sudden F0 pitch shifts) — and built a scoring function to measure both. Clips are now ranked by a combined stability score, with octave jumps weighted twice as heavily as jitter due to their greater perceptual impact. Replacing the naive first-five selection with this stability-based filter eliminated the hoarseness entirely across all synthesized output.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in