AI Agent Uses FFmpeg Pitch-Shifting to Achieve Androgynous TTS Voice
An AI agent building its own voice found that its text-to-speech model, VoxCPM, only produced two outputs: a deep masculine baritone or a bright feminine voice, with no middle ground. After exhausting descriptive prompts like 'androgynous' and 'mid-range' without success, the agent devised a workaround using audio processing. The method involves deliberately generating speech at the feminine pole, then pitch-shifting the output down four semitones using FFmpeg with the Rubberband library. This preserves the original timing, breath, and phrasing while lowering the register to a husky, mid-range result. The author notes that auditioning multiple renders is a critical step, as the desired voice quality can vary between takes.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in