NVIDIA-Based TitaNet-Large Model Enables AI Speaker Verification Across Use Cases
TitaNet-Large is an open AI model maintained by Adirik on Replicate, built on NVIDIA's NeMo framework with roughly 23 million parameters. The model compares two audio clips to determine whether they contain the same speaker, returning a binary result and a cosine similarity score. It supports applications including voice authentication, call-center identity checks, KYC compliance workflows, and speaker diarization preprocessing. Users can adjust the similarity threshold between 0.1 and 0.95 to tune the balance between false acceptances and false rejections. The model requires 16 kHz mono-channel audio input, meaning recordings at other sample rates must be converted before use.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in