glm_moe_dsa¶
Model type |
|
Class |
|
Task |
|
Source |
|
Description¶
GLM-5.2 (zai-org/GLM-5.2) causal LM: MLA + DSA + MoE.
model_type: glm_moe_dsa
DSA (top-k sparse attention via the frozen pkg.nxrt::IndexShare op)
is the default export path (config.use_dsa=True, the default).
Setting config.use_dsa=False (the --glm-full-attention CLI
feature) instead builds the plain dense-MLA DeepSeekV3TextModel
unchanged, which any Attention-op-supporting runtime – including
stock ORT – can already execute. preprocess_weights drops every DSA
indexer weight in that case since nothing in the dense graph consumes it.
Multi-token-prediction (MTP) is out of scope this cycle: see the module
docstring. preprocess_weights drops model.layers.N.* weights for
N >= num_hidden_layers (MTP) with a logged capability message rather
than silently or accidentally feeding them into the base decoder.
Usage¶
mobius build --model <MODEL_ID> output_dir/
from mobius import build
model = build("<MODEL_ID>")