Speaker diarization — working out who spoke when — has been one of those tasks that works in demos and disappoints in practice. The conventional approach chains together voice activity detection, segmentation, speaker embedding and clustering, each with its own tuning, and the whole pipeline degrades the moment two people talk over each other.
NVIDIA Nemotron 3 Diarization, released 23 September 2026, replaces that chain with one model.
The specifications
- ~100M parameters — a transformer architecture built on NVIDIA Streaming Sortformer
- Up to 8 speakers
- End-to-end: a single pass replacing segmentation-plus-embedding and its tuning
- Streaming and offline inference
- Handles overlapping speech
- Chunked processing for flexible recording lengths
- Configurable latency, and this is the interesting part:
| Mode | Latency |
|---|---|
| Ultra-low | 0.32 s |
| Very low | 0.64 s |
| Low | 1.04 s |
| Offline | 30.4 s |
DER on DIHARD III (full dataset): 12.73% at 30.4s, 13.18% at 1.04s, 13.55% at 0.32s. It ranks #1 on VoiceArena’s Diarization-Bench with a 14.72% DER.
Look at that latency/accuracy curve, because it’s the headline result hiding behind the headline number. Going from 30-second offline processing to 320-millisecond streaming costs less than one point of DER. Real-time diarization has historically meant accepting a large accuracy penalty. Here it’s almost free.
The licence, stated plainly
OpenMDW License Agreement, version 1.1. The model card says: “This model is ready for commercial or non-commercial use.”
Worth being explicit, because at least one write-up of this release described it as an evaluation-only preview. It isn’t — we checked the model card directly. Commercial use is permitted.
After a week in which the licence has been the least clear thing about nearly every model we’ve covered — Qwen-Image-2.1 research-only, Apple’s LensVLM on a research licence, Pixal3D with an unresolved MIT-versus-academic contradiction — a 100M-parameter model with a plain commercial-use statement is genuinely refreshing.
Why a music and audio site cares about diarization
Because “who spoke when” generalises to “which source is active when,” and that’s a segmentation problem that shows up everywhere in audio work:
- Interview and podcast production — auto-labelled speaker tracks, which is the single most tedious part of editing multi-person audio recorded on one mic
- Archival and oral history — making large recorded collections searchable by speaker
- Live captioning and subtitling for performance, talks and installations, where 320ms is fast enough to feel immediate
- Rehearsal and session documentation — labelling who said what across hours of recorded studio talk
- Installations that respond to multiple visitors speaking, where knowing there are three distinct voices and which is active matters more than knowing the words
That last one is the genuinely new capability for interactive work. A 100M-parameter model is small enough to run alongside other things on modest hardware, streaming at a third of a second, tracking up to eight people. An installation that knows how many people are talking and can follow turn-taking between them is now a reasonable thing to build rather than a research project.
The caveat to hold: 13–14% DER means roughly one in seven segments is mislabelled. That’s state of the art and it is not transcription-grade accuracy. Design for it being usually right, not always.
Related Reading
- nvidia/Nemotron-3-Diarization — Hugging Face
- Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization — Hugging Face blog
- NVIDIA Releases Nemotron 3 Diarization Open-Weight Speaker Model — Unite.AI
- NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model — MarkTechPost
- NVIDIA Nemotron 3 Diarization: real-time speaker labels at a cent per audio hour — Baseten