Voice Agents
Voice agents conduct spoken interactions while coordinating speech processing, model decisions and tools. They may use an audio-native model or a speech-to-text, language-model and text-to-speech pipeline, but both designs require reliable turn-taking, interruption handling and confirmation of actions that users request by voice.
What it is
A voice interaction is a streaming conversation rather than a sequence of complete text messages. The system captures audio, identifies turns, interprets the request and produces speech while tracking tool activity and what the user has actually heard. Chained pipelines expose intermediate text and allow independent component choices; audio-native approaches can preserve more speech information within one session. These architectures have different latency and control tradeoffs. Voice activity detection estimates when someone is speaking, but conversational turn completion is a separate decision. State must also distinguish generated audio from delivered audio when a user interrupts a response.
What the work involves
The practitioner chooses an audio architecture, transport and turn-detection policy based on the application. They design interruption behavior, timeouts, tool-result announcements and escalation paths. Sensitive entities such as addresses need confirmation because speech recognition can be uncertain. Evaluation includes noisy environments, accents, overlapping speech and delayed tools. Useful outputs include conversation traces aligned with audio events and a tested policy for resuming after interruption. Perceived responsiveness must be measured together with task accuracy and whether spoken confirmations match committed actions.
Illustrative example
A caller asks a booking assistant to change an appointment to Thursday afternoon. The agent repeats the proposed date and time before submitting the update. While it describes an unavailable slot, the caller interrupts with a different preference; playback stops and the new turn is interpreted using the latest booking state. A test checks that the previously spoken alternative was never committed and that the final confirmation corresponds to the appointment returned by the booking tool.
Limits and common mistakes
Fast speech output can still deliver a wrong action or talk over the user. Transcription errors, uncertain names and tool latency can break an otherwise fluent conversation. A pipeline needs explicit handling for partial input and cancelled output, not just lower average response time. Quality includes intelligibility, interruption recovery, accurate entity capture and clear handoff when speech is insufficient. Voice interfaces also need accessible alternatives for people or environments where speaking and listening are impractical.
Prerequisites
- hardAI Agent Design
A voice agent is an agent with a speech interface.
- hardAudio AI
Requires streaming speech recognition and synthesis.
Related skills
- ← is an instance of: LiveKit
Sources and further reading
- OpenAI voice agents guide
Compares audio-native, continuous voice and chained architectures and their implications for agent workflows.
Last updated: 2026-10-10