Skip to main content

load_audio

Opens an audio file and returns it as a 16 kHz mono float32 numpy array. Uses ffmpeg internally to decode the file, down-mix to mono, and resample to 16 kHz. The returned array can be passed directly to model.transcribe(), align(), or DiarizationPipeline.

Parameters

string
required
Path to the audio file. Accepts any format supported by ffmpeg, including mp3, wav, m4a, flac, ogg, opus, and aac.
integer
default:"16000"
Target sample rate for resampling. Defaults to SAMPLE_RATE (16000 Hz). All whisper(ml)x functions expect 16 kHz audio — only change this if you have a specific reason to resample to a different rate.

Returns

A NumPy array of shape (samples,) with dtype=float32. Values are normalized to the range [-1.0, 1.0].

Full signature


SAMPLE_RATE

The sample rate expected by all whisper(ml)x functions: 16000 Hz. All audio passed to transcribe(), align(), and DiarizationPipeline must be at this sample rate. load_audio() resamples to this rate automatically.

Usage

ffmpeg must be installed and available on your PATH. On macOS with Homebrew: brew install ffmpeg. load_audio raises a RuntimeError if ffmpeg is not found or if the file cannot be decoded.

MLXWhisperPipeline.transcribe

Pass audio arrays directly to transcription.

Schema & Types

TypedDict definitions for transcription output.