Tencent's Gander AI Splits Brain Architecture to Handle Voice Chat and Background Tasks Simultaneously

Tencent's Hunyuan Speech Team released an open-weight AI research project called Gander in early September, designed to hold real-time voice conversations while executing background tasks concurrently. The system uses a dual-component architecture inspired by the human brain: a 'Cerebellum' module handles live voice and visual input, while a 'Brain' module processes heavier tasks like web searches or code execution in the background. Model weights are publicly available on Hugging Face under an Apache 2.0 license, built on top of the open MiniCPM-o 4.5 base model. However, running the full system requires three separate GPUs and approximately 19 GB of VRAM in full precision, making it impractical for most general users at this stage. In benchmark tests, Gander achieved around 40% task-completion accuracy on overlapping-speech evaluations, compared to roughly 60% for closed systems like GPT-Realtime, positioning it as a research milestone rather than a consumer-ready product.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in