moonshine_streaming¶
Model type |
|
Class |
|
Task |
|
Source |
|
Description¶
Moonshine Streaming encoder-decoder model for low-latency ASR.
Replicates Hugging Face MoonshineStreamingForConditionalGeneration: a
causal framing/convolution front end feeding a windowed (bounded-lookahead)
transformer encoder, and a cached Moonshine decoder that cross-attends to the
position-adapted encoder context.
Caller contract¶
The export contains an encoder and a cached decoder. It supports
FP16 CUDA builds (execution_provider="cuda") and CPU builds.
The encoder consumes a rolling raw-audio window, not precomputed features:
Tensor |
Type |
Shape |
Meaning |
|---|---|---|---|
|
graph dtype ( |
|
Raw 16 kHz waveform. The sample dimension must be a multiple of 80: the model frames 80 samples (5 ms) at a time. |
|
|
|
|
|
graph dtype |
|
Encoder output for the decoder. |
|
|
|
Downsampled validity mask for encoder frames. |
The causal convolution front end downsamples four 5-ms input frames into each
encoder frame. Per-layer sliding_windows gives encoder attention bounded
right lookahead, so frames near a window’s right edge can change after more
audio arrives. The pinned tiny checkpoint has four layers with four
right-lookahead encoder frames, giving a maximum composed 16-frame (320-ms)
lookahead. The export has no stable-frame-count output; callers must retain
that tail when displaying a provisional result.
This is a streaming architecture, not a stateful streaming inference API.
Neither this export nor Hugging Face’s
MoonshineStreamingForConditionalGeneration provides encoder cache,
encoder offset, predict, commit, or finalize inputs/outputs.
Each encoder invocation is independent, so callers cannot faithfully commit
newly stable tokens by chaining ONNX encoder state. The supported protocol is
to retain the raw waveform, periodically re-run the encoder and regular
seq2seq decoder for an uncommitted hypothesis, and publish the final
transcript after end-of-stream. At end-of-stream, pad the final raw window to
an 80-sample boundary, set its padding-mask entries to zero, re-run on the
complete buffered waveform, and decode normally; there is no separate flush
operation or finalization token.
To decode, pass decoder_input_ids, the current encoder outputs, and
position_ids, with zero-length past_key_values.*.{key,value} tensors
for the first token. Feed each present.*.{key,value} output back as its
corresponding past_key_values.*.{key,value} input for later tokens while
keeping the encoder outputs and mask fixed. If an application displays
partial hypotheses, it must compare and de-duplicate those ordinary decoded
token sequences at the application layer; they are not model commit events.
Usage¶
mobius build --model <MODEL_ID> output_dir/
from mobius import build
model = build("<MODEL_ID>")