cosmos3_edge

Model type

cosmos3_edge

Class

Cosmos3EdgeVLModel

Task

cosmos3-edge-vl

Source

models/cosmos.py

Description

NVIDIA Cosmos3-Edge vision-language model (3-model split).

model_type: cosmos3_edge / Cosmos3EdgeForConditionalGeneration.

Builds three ONNX models for onnxruntime-genai deployment:

  • decoder: squared-ReLU GQA text reasoner taking inputs_embeds.

  • vision_encoder: SigLIP vision tower + pixel-shuffle merger projector.

  • embedding: token embedding + image feature fusion at image_token_id (19).

HuggingFace weight layout (single checkpoint, no language_model. prefix):

  • model.visual.* → vision tower (SigLIP)

  • model.projector.* → merger projector

  • embed_tokens.weight → embedding

  • layers.* / norm.weight / lm_head.weight → decoder

  • *k_norm_und_for_gen* → dropped (generator-tower artifact; see the module docstring)

.. note:: NVIDIA does not publish modeling code for cosmos3_edge (it is not in transformers and the repo ships no remote-code module), so the exact pixel-shuffle ordering and numerical parity are unverifiable. This build is validated at graph-construction (L1) confidence only.

Usage

mobius build --model <MODEL_ID> output_dir/
from mobius import build

model = build("<MODEL_ID>")