Setup & Installation
What This Skill Does
Python SDK for building real-time voice AI applications over WebSocket connections with Azure AI. Handles bidirectional audio streaming, voice activity detection, function calling, and transcription in a single async interface. Connects to models like GPT-4o Realtime for low-latency speech-to-speech interactions.
Manually managing WebSocket audio streams, VAD, interruption handling, and event routing would require significant custom infrastructure — this SDK wraps all of that into a typed async Python interface.
When to use it
- Building a voice assistant that detects speech pauses and responds automatically using Server VAD
- Adding real-time voice input and audio output to a Python chatbot
- Routing spoken function calls to backend tools during a live voice session
- Streaming PCM audio at different sample rates for telephony versus high-quality voice apps
- Transcribing user speech in real time while simultaneously generating an audio response