VibeVoiceForASRStreamingTraining

Model type

VibeVoiceForASRStreamingTraining

Class

VibeVoiceASRStreamingForConditionalGeneration

Task

vibevoice-asr-streaming

Source

models/vibevoice.py

Description

Streaming ASR package for VibeVoiceForASRStreamingTraining checkpoints.

This shared implementation supports the pinned Microsoft 1.5B and 7B streaming-ASR checkpoints. Unlike native offline ASR, it executes two cached causal tokenizers and their connectors in one audio_encoder stage. The host supplies speech masks and reproducible acoustic noise, inserts flattened speech embeddings at placeholder positions, then caches Qwen2 decoding.

Usage

mobius build --model <MODEL_ID> output_dir/
from mobius import build

model = build("<MODEL_ID>")