glm_moe_dsa

Model type

glm_moe_dsa

Class

GlmMoeDsaCausalLMModel

Task

text-generation

Source

models/glm_moe_dsa.py

Description

GLM-5.2 (zai-org/GLM-5.2) causal LM: MLA + DSA + MoE.

model_type: glm_moe_dsa

DSA (top-k sparse attention via the frozen pkg.nxrt::IndexShare op) is the default export path (config.use_dsa=True, the default). Setting config.use_dsa=False (the --glm-full-attention CLI feature) instead builds the plain dense-MLA DeepSeekV3TextModel unchanged, which any Attention-op-supporting runtime – including stock ORT – can already execute. preprocess_weights drops every DSA indexer weight in that case since nothing in the dense graph consumes it.

Multi-token-prediction (MTP) is out of scope this cycle: see the module docstring. preprocess_weights drops model.layers.N.* weights for N >= num_hidden_layers (MTP) with a logged capability message rather than silently or accidentally feeding them into the base decoder.

Usage

mobius build --model <MODEL_ID> output_dir/
from mobius import build

model = build("<MODEL_ID>")