Voice Activity Detection
NeMo
Safetensors
GGUF
Transformers
nemotron3_diarization
audio-frame-classification
speaker-diarization
streaming-sortformer
speaker-tagging
Instructions to use nvidia/Nemotron-3-Diarization with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use nvidia/Nemotron-3-Diarization with NeMo:
# tag did not correspond to a valid NeMo domain.
- Transformers
How to use nvidia/Nemotron-3-Diarization with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForAudioFrameClassification processor = AutoProcessor.from_pretrained("nvidia/Nemotron-3-Diarization") model = AutoModelForAudioFrameClassification.from_pretrained("nvidia/Nemotron-3-Diarization", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download explainability.md from nvidia/Nemotron-3-Diarization: direct link, hf CLI and curl.
- Browser
- Download file 2.33 kB
-
https://huggingface.co/nvidia/Nemotron-3-Diarization/resolve/main/explainability.md
- Command line
-
hf download hf://nvidia/Nemotron-3-Diarization/explainability.md
-
curl -L -o explainability.md https://huggingface.co/nvidia/Nemotron-3-Diarization/resolve/main/explainability.md
2.33 kB
| Field | Response |
|---|---|
| Intended Task/Domain | Speaker Diarization (determining who spoke when in conversational audio), Speaker Tagging in Speech Recognition Systems |
| Model Type | Transformer encoder |
| Intended Users | Developers, researchers and agents working on conversational AI systems that require generic speaker labels and timestamps. |
| Output | A [T, 8] floating-point tensor of per-speaker activity probabilities. The eight channels are ordered by speaker arrival time. The default output stride is 10 ms and can be configured to a multiple of 10 ms. |
| Describe how the model works | The model stacks 10 ms Mel-spectrogram features by a factor of eight, processes the resulting 80 ms frames with a Transformer encoder, and upsamples predictions back to 10 ms resolution with a Conv1D layer. During streaming inference, the Arrival-Order Speaker Cache and FIFO queue preserve useful speaker context across chunks. Speaker activity probabilities can be postprocessed into generic speaker labels with start and end timestamps. |
| Name the adversely impacted groups this has been tested to deliver comparable outcomes regardless of | Not Applicable |
| Technical Limitations & Mitigation | This model supports up to eight speakers; if the audio contains more than eight speakers, some will inevitably be missed or misattributed. Performance may degrade on very long recordings and in challenging acoustic conditions, such as high background noise or severe reverberation. Latency profile should be chosen carefully for a downstream application, as lowering latency generally decreases both accuracy and speed of the model. |
| Verified to have met prescribed NVIDIA quality standards | Yes |
| Performance Metrics | Diarization Error Rate (DER), Speaker Counting Accuracy (SCA), Speaker Counting Mean Absolute Error (MAE), and Real-Time Factor speedup (RTFx). |
| Potential Known Risks | Diarization errors can assign speech to the wrong generic speaker channel, miss speech, introduce false alarms, or place boundaries incorrectly, particularly under overlap, noise, far-field conditions, domain shift, or unsupported speaker counts. Downstream systems should retain uncertainty and be tested for the intended use case. |
| Licensing | Use of this model is governed by the OpenMDW License Agreement, version 1.1. |