Skip to main content
whisper(ml)x can write transcription results to several file formats. The default (all) produces every format simultaneously. Use --output_format to select a specific one.

Available formats

CLI usage

Output files are named after the input file. Transcribing audio.mp3 produces audio.srt, audio.json, etc. in the output directory.

Format details

SRT

Standard subtitle format supported by most video players and editing tools:
When diarization is enabled, speaker labels are prepended to each cue:

VTT

WebVTT format for embedding in HTML5 video via the <track> element:

TXT

Plain text, one segment per line. Speaker labels are included when diarization data is present:
With diarization:

TSV

Tab-separated values with a header row. Timestamps are in integer milliseconds:

JSON

Full output including all segment data, word-level timestamps (when alignment ran), and speaker labels (when diarization ran):
The speaker_embeddings key is only present when --speaker_embeddings was passed.

AUD

Audacity label file format. Timestamps are in seconds (not milliseconds). The file can be imported into Audacity via File → Import → Labels:
With diarization, the speaker is embedded in double brackets:
The .aud format is not included in all. It must be requested explicitly with --output_format aud.

Subtitle formatting options

These options apply to SRT and VTT output and require alignment (they are incompatible with --no_align):

Line wrapping example

Word highlighting

--highlight_words uses HTML underline tags (<u>word</u>) to mark the active word in each cue:
--max_line_width, --max_line_count, and --highlight_words require word-level timestamps from alignment. Using them with --no_align will cause an error.

Segment resolution

--segment_resolution is accepted as a CLI flag (default: sentence) for compatibility, but does not currently affect subtitle output directly.

Output directory

Use --output_dir (short: -o) to set where files are written. The default is the current working directory:

Writing output from Python

Use get_writer to write results programmatically:
Pass "all" to write every format (excluding aud) in one call:

Next steps

Transcription

Transcribe audio files to get a result to write.

Alignment

Add word timestamps to unlock subtitle formatting options.

Diarization

Add speaker labels to SRT, VTT, TXT, and JSON output.

CLI reference

Full reference for all CLI flags.