cosmos3_omni¶
Model type |
|
Class |
|
Task |
|
Source |
|
Description¶
Cosmos3-Omni understanding tower (Qwen3-VL 3-model split).
Identical graph to :class:Qwen3VL3ModelCausalLMModel; only weight
preprocessing differs because the unified Cosmos3 checkpoint uses NVIDIA’s
native parameter names and carries the extra Generator / Sound / Action
towers.
.. note::
DeepStack is preserved. The Qwen3-VL 3-model split packs the
intermediate maps (from deepstack_visual_indexes — [8, 16, 24]
for Cosmos3) into image_features. The embedding model scatters and
flattens them into ORT GenAI’s rank-3 per_layer_inputs contract,
and the decoder restores and injects one map into each of its first
D layers — matching HuggingFace’s per-layer injection. The 18
deepstack_merger_list.* weights are therefore exported and routed
to vision_encoder.
Usage¶
mobius build --model <MODEL_ID> output_dir/
from mobius import build
model = build("<MODEL_ID>")