Vowel estimator accuracy skewed by audio order and duration, developer finds
A developer building a browser-based vowel estimator for VRM avatar lip-syncing discovered that evaluation results were being distorted by the measurement method itself rather than true estimator performance. Because the estimator tracks a long-term average of frequency-band levels that updates with each audio input, the same sound produces different feature values depending on what was heard before it. Evaluation audio was synthesized using Style-Bert-VITS2 across three speakers and five Japanese vowels, then mechanically screened for length, RMS, peak, voicing rate, and formant quality before use. Pitfalls included misleading peak normalization flagged as clipping and LPC-based formant estimation picking up harmonics for high-pitched speakers, prompting a direct spectral analysis instead. A leave-one-speaker-out validation scheme was also applied to ensure results were not artificially inflated by speaker overlap between template design and evaluation data.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in