nemotron3_diarization

Model type

nemotron3_diarization

Class

Nemotron3DiarizationModel

Task

diarization

Source

models/nemotron3_diarization.py

Description

Streaming Sortformer speaker-diarization model (HuggingFace port).

Replicates HuggingFace’s Nemotron3DiarizationForAudioFrameClassification (e.g. nvidia/Nemotron-3-Diarization): a bidirectional, partial-RoPE Transformer audio encoder followed by a sub-pixel upsampler and a speaker sigmoid head. Consumes mel-spectrogram features and returns per-frame speaker-activity probabilities in [0, 1].

Two forwards are exported (see tasks/_diarization.py and tasks/_diarization_streaming.py):

  • meth:

    forward — the offline whole-recording pass (diarization task), chunked via an ONNX Loop exactly like HuggingFace’s offline forward (any recording length, not just single-chunk ones).

  • meth:

    forward_streaming — the streaming, per-chunk pass (diarization-streaming task): consumes one chunk of audio plus a few look-ahead frames, together with the previous step’s Arrival-Order Speaker Cache (AOSC) + FIFO queue state (fixed-size buffers and scalar occupancy counters, all graph inputs/outputs), and returns this chunk’s speaker probabilities plus the updated cache state. Repeated calls reproduce HuggingFace’s streaming Nemotron3DiarizationSpeakerCache bookkeeping exactly, including its top-k score-based compression.

    Limitation: unlike HuggingFace, the streaming export does not accept a padding attention_mask — every step’s input window is assumed fully valid (no silence padding within a chunk). This holds for all but possibly the very last chunk of a recording, matching the precision needed for real-time streaming use.

Usage

mobius build --model <MODEL_ID> output_dir/
from mobius import build

model = build("<MODEL_ID>")