all) produces every format simultaneously. Use --output_format to select a specific one.
Available formats
CLI usage
audio.mp3 produces audio.srt, audio.json, etc. in the output directory.
Format details
SRT
Standard subtitle format supported by most video players and editing tools:VTT
WebVTT format for embedding in HTML5 video via the<track> element:
TXT
Plain text, one segment per line. Speaker labels are included when diarization data is present:TSV
Tab-separated values with a header row. Timestamps are in integer milliseconds:JSON
Full output including all segment data, word-level timestamps (when alignment ran), and speaker labels (when diarization ran):speaker_embeddings key is only present when --speaker_embeddings was passed.
AUD
Audacity label file format. Timestamps are in seconds (not milliseconds). The file can be imported into Audacity via File → Import → Labels:The
.aud format is not included in all. It must be requested explicitly with --output_format aud.Subtitle formatting options
These options apply to SRT and VTT output and require alignment (they are incompatible with--no_align):
Line wrapping example
Word highlighting
--highlight_words uses HTML underline tags (<u>word</u>) to mark the active word in each cue:
Segment resolution
--segment_resolution is accepted as a CLI flag (default: sentence) for compatibility, but does not currently affect subtitle output directly.
Output directory
Use--output_dir (short: -o) to set where files are written. The default is the current working directory:
Writing output from Python
Useget_writer to write results programmatically:
"all" to write every format (excluding aud) in one call:
Next steps
Transcription
Transcribe audio files to get a result to write.
Alignment
Add word timestamps to unlock subtitle formatting options.
Diarization
Add speaker labels to SRT, VTT, TXT, and JSON output.
CLI reference
Full reference for all CLI flags.