cosmos3_edge¶
Model type |
|
Class |
|
Task |
|
Source |
|
Description¶
NVIDIA Cosmos3-Edge vision-language model (3-model split).
model_type: cosmos3_edge / Cosmos3EdgeForConditionalGeneration.
Builds three ONNX models for onnxruntime-genai deployment:
decoder: squared-ReLU GQA text reasoner taking
inputs_embeds.vision_encoder: SigLIP vision tower + pixel-shuffle merger projector.
embedding: token embedding + image feature fusion at
image_token_id(19).
HuggingFace weight layout (single checkpoint, no language_model.
prefix):
model.visual.*→ vision tower (SigLIP)model.projector.*→ merger projectorembed_tokens.weight→ embeddinglayers.*/norm.weight/lm_head.weight→ decoder*k_norm_und_for_gen*→ dropped (generator-tower artifact; see the module docstring)
.. note::
NVIDIA does not publish modeling code for cosmos3_edge (it is not in
transformers and the repo ships no remote-code module), so the exact
pixel-shuffle ordering and numerical parity are unverifiable. This build
is validated at graph-construction (L1) confidence only.
Usage¶
mobius build --model <MODEL_ID> output_dir/
from mobius import build
model = build("<MODEL_ID>")