Setup & Installation
What This Skill Does
Transcribes audio files to text using OpenAI's transcription models. Supports plain text output and speaker diarization with known-speaker hints for labeled transcripts. Works via a bundled Python CLI with configurable chunking and output formats.
Using the bundled CLI with auto-chunking handles long files and multi-speaker labeling in one command, without manually splitting audio or post-processing speaker attribution.
When to use it
- Transcribing a recorded interview with labeled speaker turns
- Extracting text from a meeting recording saved as an .m4a file
- Converting a voice memo to a plain text file for editing
- Running diarization on a podcast episode to separate two known hosts
- Batch transcribing multiple audio files into separate output directories