Whisper transcribes speech extremely well and is completely indifferent to who is speaking. Feed it a two-person interview and you get a wall of accurate text with no indication that the speaker ever changed.
Getting “who spoke when” is a separate task — speaker diarization — handled by a separate model. This is how you join them, running locally, with no API bill.
What you’re building
Three stages:
- Transcription — Whisper (or WhisperX) → text with timestamps
- Diarization — pyannote.audio → time ranges labelled
SPEAKER_00,SPEAKER_01, … - Merge — assign each transcript segment to whichever speaker was talking during it
Neither model knows about the other. Stage 3 is all yours, and it’s where most of the quality lives.
Setup
pip install openai-whisper pyannote.audio torch
# or, strongly recommended:
pip install whisperx
Use WhisperX if you can. It wraps Whisper with forced alignment to produce word-level timestamps rather than segment-level ones, and it has diarization integration built in. Plain Whisper gives you timestamps per segment, which are often several seconds long and frequently span a speaker change — and a segment that spans a speaker change cannot be assigned to one speaker correctly.
pyannote’s pretrained pipelines are gated on Hugging Face: accept the model conditions on the model page and authenticate with a token, or it will fail with a permissions error that doesn’t obviously say that’s the problem.
The short version, with WhisperX
import whisperx
device = "cuda" # "cpu" works, slowly
audio_file = "interview.wav"
# 1. transcribe
model = whisperx.load_model("large-v3", device, compute_type="float16")
audio = whisperx.load_audio(audio_file)
result = model.transcribe(audio, batch_size=16)
# 2. align -> word-level timestamps
model_a, metadata = whisperx.load_align_model(
language_code=result["language"], device=device
)
result = whisperx.align(
result["segments"], model_a, metadata, audio, device
)
# 3. diarize and assign
diarize_model = whisperx.DiarizationPipeline(
use_auth_token="hf_...", device=device
)
diarize_segments = diarize_model(audio, min_speakers=2, max_speakers=4)
result = whisperx.assign_word_speakers(diarize_segments, result)
for seg in result["segments"]:
print(f"[{seg.get('speaker','?')}] {seg['text'].strip()}")
Pass min_speakers and max_speakers whenever you know them. Diarization has to guess how many people are present, and guessing is where it goes most wrong — usually by splitting one person into two when their voice changes across a recording. If you know it’s a two-person interview, say so.
Doing the merge yourself
If you’re using plain Whisper and pyannote separately, the merge is a time-overlap problem. The naive version — assign by segment midpoint — works acceptably:
def speaker_at(t, diarization):
for turn, _, speaker in diarization.itertracks(yield_label=True):
if turn.start <= t <= turn.end:
return speaker
return None
for seg in whisper_result["segments"]:
mid = (seg["start"] + seg["end"]) / 2
seg["speaker"] = speaker_at(mid, diarization)
The better version assigns per word and then groups runs of consecutive words by speaker. That’s what WhisperX does, and it’s why word-level timestamps matter: it handles the case where a Whisper segment genuinely contains two speakers, which the midpoint method silently gets half wrong.
What accuracy to expect
Set expectations correctly or you’ll be disappointed for the wrong reasons.
State-of-the-art diarization sits around 12–15% Diarization Error Rate on hard benchmarks — NVIDIA’s new Nemotron 3 Diarization tops VoiceArena’s leaderboard at 14.72%. Roughly one segment in seven is mislabelled, and that’s the best available.
So: transcription is usually production-ready; speaker labels are a good first pass that a human should check. Do not ship auto-labelled speaker attributions as fact in anything that matters.
What actually breaks it
- Overlapping speech. Two people talking at once is the hardest case and the most common in real conversation. Expect errors there specifically.
- Similar voices. Two speakers of the same gender with similar pitch and accent will get merged. Nothing fixes this at the model level.
- One person on multiple mics, or varying distance. The same voice recorded differently across a session can be split into two speakers.
- Short interjections. “Mm-hm”, “right”, “yeah” are frequently dropped or assigned to the wrong person.
- Music and noise. Diarization degrades badly with background music. Separate stems first if you can.
- Speaker labels are not stable across files.
SPEAKER_00in one recording is not the same person asSPEAKER_00in the next. Mapping labels to actual names is a manual step, every time.
Where to go next
Once the pipeline runs: faster-whisper for a substantial speed-up via CTranslate2, VAD preprocessing to strip silence before transcription, and NVIDIA’s Nemotron 3 Diarization as a pyannote alternative worth benchmarking — it’s ~100M parameters, streams at 320ms latency, handles up to eight speakers, and its OpenMDW 1.1 licence permits commercial use, which pyannote’s pretrained pipelines complicate.
For creative use, the interesting output isn’t the transcript. It’s having time-aligned, speaker-attributed text for hours of material — which makes an archive searchable, a documentary assemblable, and an oral history navigable by who was talking.