NVIDIA on September 23 released Nemotron 3 Diarization, an open-weight speaker diarization model with roughly 100 million parameters that identifies who spoke and when in multi-speaker recordings. On the VoiceArena benchmark, it posts a 14.72% diarization error rate — about 24% better than the previous open-source leader.
Compact, Fast, And Fully Open
The model runs on a single consumer GPU and is available through Hugging Face Transformers and the Dell Enterprise Hub under NVIDIA's open-weight license. NVIDIA is pitching it as a drop-in front end for call-center analytics, meeting summarization, media captioning and any pipeline that already runs on Nemotron speech-to-text.
Where It Slots In The Nemotron Stack
Nemotron 3 Diarization joins a growing family of open Nemotron models NVIDIA has shipped in 2026, including Nemotron 3.5 Lightning and NeMo Switchyard, the Nemotron-3-Ultra-CC competitive-coding model, and integrations powering Salesforce's Koa reasoning stack.
Why It Matters
Open-weight speaker labeling has lagged proprietary APIs from Google, Otter and AssemblyAI. Nemotron 3 Diarization narrows that gap for enterprises that need on-prem transcription, media companies handling copyrighted content, and researchers building agent stacks that need to know which participant is speaking.
Reporting based on coverage from NVIDIA, Hugging Face and industry announcements on September 23, 2026.
