nemotron3_diarization¶
Model type |
|
Class |
|
Task |
|
Source |
|
Description¶
Streaming Sortformer speaker-diarization model (HuggingFace port).
Replicates HuggingFace’s Nemotron3DiarizationForAudioFrameClassification
(e.g. nvidia/Nemotron-3-Diarization): a bidirectional, partial-RoPE
Transformer audio encoder followed by a sub-pixel upsampler and a speaker
sigmoid head. Consumes mel-spectrogram features and returns per-frame
speaker-activity probabilities in [0, 1].
Two forwards are exported (see tasks/_diarization.py and
tasks/_diarization_streaming.py):
- meth:
forward— the offline whole-recording pass (diarizationtask), chunked via an ONNXLoopexactly like HuggingFace’s offline forward (any recording length, not just single-chunk ones).
- meth:
forward_streaming— the streaming, per-chunk pass (diarization-streamingtask): consumes one chunk of audio plus a few look-ahead frames, together with the previous step’s Arrival-Order Speaker Cache (AOSC) + FIFO queue state (fixed-size buffers and scalar occupancy counters, all graph inputs/outputs), and returns this chunk’s speaker probabilities plus the updated cache state. Repeated calls reproduce HuggingFace’s streamingNemotron3DiarizationSpeakerCachebookkeeping exactly, including its top-k score-based compression.
Limitation: unlike HuggingFace, the streaming export does not accept a padding
attention_mask— every step’s input window is assumed fully valid (no silence padding within a chunk). This holds for all but possibly the very last chunk of a recording, matching the precision needed for real-time streaming use.
Usage¶
mobius build --model <MODEL_ID> output_dir/
from mobius import build
model = build("<MODEL_ID>")