Real-time voice agents need clear consent and handoff design when simultaneous listening and speaking turns AI into a continuous work interface
Source: TestingCatalog
TLDR IT highlighted a hidden early-access listing for Microsoft's MAI Realtime model. The reported design supports simultaneous listening and speaking, multilingual conversations, configurable turn-taking, two voices, and interruptions, though it has not been announced for broad release. The operational change is a move away from the request-and-wait rhythm of conventional assistants.
Why this matters: Treat a continuous voice agent as a live work surface, not merely a speech feature. Establish when recording or transcription occurs, how a user interrupts or escalates, which actions remain draft-only, and how identity, sensitive data, and audit evidence carry across a spoken workflow.
Read the report