Researchers Propose 8-Axis Perceptual Slider System for Fine-Tuning Synthetic Voices
A technique published on DEV Community describes a method to blend speaker embeddings from multiple anchor voices using eight perceptual sliders — covering age, pitch, huskiness, clarity, warmth, roughness, gender, and build. The approach addresses a core limitation in voice conversion and TTS models, which typically accept only a single reference audio file to define a target voice, making gradual adjustments difficult. Each perceptual axis is defined as a weighted combination of measurable acoustic features such as F0, formants, HNR, shimmer, and jitter, extracted from pre-recorded anchor speaker audio. User slider values are converted into target z-scores, and softmax-weighted distances to each anchor determine how their embeddings are blended into a final voice vector. The system also distinguishes between reliably measured axes and those requiring human labels, excluding uncalibrated axes from computation to preserve overall reliability.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in