SShortSingh.
Back to feed

Full-Duplex Voice AI: Why Interruption Handling Reshapes the Entire Pipeline

0
·1 views

Unlike traditional turn-based voice assistants, full-duplex systems allow both the user and the AI to speak simultaneously, requiring the assistant to stop instantly when interrupted. Achieving this demands careful engineering: playback and microphone streams must be kept separate, echo cancellation must be precise, and any pending response should be discarded the moment a user speaks. Turn detection relies on a small dedicated model running over the audio stream, tuned on real recordings rather than scripted demos, to accurately sense when a speaker has paused versus finished. Latency budgeting is critical, with each stage — turn detection, retrieval, first token, and synthesis — needing explicit time allocations to keep responses feeling responsive. Developers are advised to test interruptions at the first word, mid-sentence, and during tool calls, and to use browser-based playgrounds to evaluate model behaviour before committing to a transport layer.

Read the full story at DEV Community

This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)

Log in to join the discussion and vote.

Log in

Related stories

0
ProgrammingDEV Community ·

Developer Builds AI Incident Response Agent That Avoids Overconfident Fixes

A developer has built an open-source AI incident response tool called RecallOps, designed to assist engineers during on-call outages by referencing past incidents stored in memory. The agent uses Hindsight, an open-source memory system, paired with OpenAI's GPT-OSS-120B model served via Groq at a low temperature setting to prioritize accuracy over creativity. A core design principle is that retrieved past incidents are treated as evidence for comparison, not as ready-made solutions, preventing the model from blindly recommending fixes that may not match the current root cause. The prompt explicitly instructs the model to identify similarities and differences between past and present incidents before suggesting any action, and to flag uncertainty when evidence is insufficient. This approach was built to guard against a specific failure mode where similar-looking symptoms lead to wrong fixes that can escalate a minor outage into a larger one.

0
ProgrammingDEV Community ·

Developer Builds Personal AI Running Coach by Feeding It Full Training History

A software developer found that using ChatGPT with individual workout screenshots was ineffective because isolated data points lack the context needed for meaningful coaching advice. To solve this, he built a personal web portal called running-portal that syncs workout data from Mi Fitness, stores historical runs, and automatically assembles structured prompts for an AI model after each session. The portal feeds the AI a formatted context block containing the current run's metrics alongside recent workout history, enabling trend-based analysis rather than one-off observations. The project relies on an open-source sync module, Mi-Fitness-Sync, which handles Xiaomi's private API and data parsing. Because the integration depends on undocumented internal APIs, the developer treats the tool as a personal utility rather than a publicly deployable service.

0
ProgrammingDEV Community ·

Cisco Firewall Management Center Flaws Exploited in Wild, One Scores Perfect 10

Cisco Talos disclosed on 9 September 2026 that two vulnerabilities in the web interface of Cisco Secure Firewall Management Center were being actively exploited by three separate threat clusters. The more critical flaw, CVE-2026-20079, carries a maximum CVSS score of 10.0 and allows a pre-authentication bypass, enabling attackers to execute scripts and gain root access via a crafted HTTP request. A second vulnerability, CVE-2026-20316, scores a relatively low 5.3 but was chained with the first flaw to serve as initial access in at least one ransomware attack. Although CVE-2026-20079 was patched in March 2026, Cisco updated its advisory in September after learning in August that exploitation had occurred in the wild. Because FMC serves as the central console for managing entire Cisco firewall fleets, a compromised instance can allow attackers to alter firewall rules, erase logs, and push malicious configurations across all managed devices.

0
ProgrammingDEV Community ·

How to Build AI Features Before Your Preferred Model Is Available

When a highly anticipated AI model is announced but not yet publicly available, developers face a dilemma: wait and stall, or build on a lesser model and refactor later. A practical approach treats the model as a swappable dependency, abstracting all model-specific logic into a single function so that switching models requires only a configuration change. Developers are advised to build a standardised test set of real, edge-case inputs to objectively evaluate any new model against the current one upon release. Per-model feature flags allow gradual traffic rollouts, reducing the risk of a full commitment before performance is confirmed. Transparent UI communication is also stressed — interfaces should clearly indicate which model is actually being used rather than implying unreleased capabilities are already live.