Setup & Installation
What This Skill Does
Transcribes short audio files (up to 60 seconds) using Azure's Speech to Text REST API in Python. No SDK required — just HTTP requests with an API key and region. Supports WAV and OGG formats, multiple languages, and both simple and detailed response formats.
Using the REST API directly avoids the Speech SDK dependency and setup overhead, making it practical for lightweight scripts, containers, or environments where installing the full SDK is impractical.
When to use it
- Transcribing voice memos recorded on a phone before storing them in a database
- Converting short customer support call clips to text for ticket logging
- Adding captions to brief video segments under one minute
- Parsing spoken commands from a voice UI without installing the Speech SDK
- Extracting text from audio collected in a lightweight Python script or serverless function