load_audio
model.transcribe(), align(), or DiarizationPipeline.
Parameters
string
required
Path to the audio file. Accepts any format supported by ffmpeg, including mp3, wav, m4a, flac, ogg, opus, and aac.
integer
default:"16000"
Target sample rate for resampling. Defaults to
SAMPLE_RATE (16000 Hz). All whisper(ml)x functions expect 16 kHz audio — only change this if you have a specific reason to resample to a different rate.Returns
A NumPy array of shape(samples,) with dtype=float32. Values are normalized to the range [-1.0, 1.0].
Full signature
SAMPLE_RATE
transcribe(), align(), and DiarizationPipeline must be at this sample rate. load_audio() resamples to this rate automatically.
Usage
ffmpeg must be installed and available on your
PATH. On macOS with Homebrew: brew install ffmpeg. load_audio raises a RuntimeError if ffmpeg is not found or if the file cannot be decoded.Related
MLXWhisperPipeline.transcribe
Pass audio arrays directly to transcription.
Schema & Types
TypedDict definitions for transcription output.