phimoe_gguf¶
Model type |
|
Class |
|
Task |
|
Source |
|
Description¶
PhiMoE graph matching llama.cpp’s GGUF inference semantics.
The HuggingFace model uses SparseMixer routing, while the pinned llama.cpp loader uses full softmax followed by normalized top-k routing. GGUF import therefore uses this internal graph without changing native HF builds.
Usage¶
mobius build --model <MODEL_ID> output_dir/
from mobius import build
model = build("<MODEL_ID>")