cosmos3_omni

Model type

cosmos3_omni

Class

Cosmos3OmniReasonerModel

Task

qwen-vl

Source

models/cosmos3_omni.py

Description

Cosmos3-Omni understanding tower (Qwen3-VL 3-model split).

Identical graph to :class:Qwen3VL3ModelCausalLMModel; only weight preprocessing differs because the unified Cosmos3 checkpoint uses NVIDIA’s native parameter names and carries the extra Generator / Sound / Action towers.

.. note:: DeepStack is preserved. The Qwen3-VL 3-model split packs the intermediate maps (from deepstack_visual_indexes — [8, 16, 24] for Cosmos3) into image_features. The embedding model scatters and flattens them into ORT GenAI’s rank-3 per_layer_inputs contract, and the decoder restores and injects one map into each of its first D layers — matching HuggingFace’s per-layer injection. The 18 deepstack_merger_list.* weights are therefore exported and routed to vision_encoder.

Usage

mobius build --model <MODEL_ID> output_dir/
from mobius import build

model = build("<MODEL_ID>")