Supported Models

mobius supports 298 registered model types.

Text Generation

Standard autoregressive language models (CausalLM).

model_type

Class

Task

DFlashDraftModel

DFlashDraftModel

dflash-draft

Eagle3DraftModel

Eagle3DraftModel

eagle3-draft

Eagle3LlamaForCausalLM

Eagle3DraftModel

eagle3-draft

Eagle3Speculator

Eagle3DraftModel

eagle3-draft

Gemma4AssistantForCausalLM

Gemma4AssistantCausalLMModel

gemma4-assistant

Gemma4UnifiedAssistantForCausalLM

Gemma4AssistantCausalLMModel

gemma4-assistant

LlamaForCausalLMEagle3

Eagle3DraftModel

eagle3-draft

Qwen35MtpModel

Qwen35MtpModel

qwen35-mtp

apertus

ApertusCausalLMModel

text-generation

arcee

ArceeCausalLMModel

text-generation

baichuan

CausalLMModel

text-generation

bloom

BloomCausalLMModel

text-generation

chatglm

ChatGLMCausalLMModel

text-generation

code_llama

CausalLMModel

text-generation

codegen

CodeGenCausalLMModel

text-generation

codegen2

CausalLMModel

text-generation

cohere

CohereCausalLMModel

text-generation

cohere2

CohereCausalLMModel

text-generation

command_r

CausalLMModel

text-generation

csm

CausalLMModel

text-generation

diffllama

DiffLlamaCausalLMModel

text-generation

doge

DogeCausalLMModel

text-generation

ernie4_5

ErnieCausalLMModel

text-generation

evolla

CausalLMModel

text-generation

exaone

CausalLMModel

text-generation

exaone4

ExaOne4CausalLMModel

text-generation

falcon

FalconCausalLMModel

text-generation

falcon_h1

FalconCausalLMModel

text-generation

falcon_mamba

MambaCausalLMModel

ssm-text-generation

gemma

GemmaCausalLMModel

text-generation

gemma2

Gemma2CausalLMModel

text-generation

gemma3_text

Gemma3CausalLMModel

text-generation

gemma3n

Gemma3nCausalLMModel

text-generation

gemma3n_text

Gemma3nCausalLMModel

text-generation

gemma4_assistant

Gemma4AssistantCausalLMModel

gemma4-assistant

gemma4_unified_assistant

Gemma4AssistantCausalLMModel

gemma4-assistant

glm

GlmCausalLMModel

text-generation

glm4

Glm4CausalLMModel

text-generation

glm4v_text

Glm4CausalLMModel

text-generation

gpt_neox

GPTNeoXCausalLMModel

text-generation

gpt_neox_japanese

GPTNeoXJapaneseCausalLMModel

text-generation

gptj

GPTJCausalLMModel

text-generation

granite

GraniteCausalLMModel

text-generation

helium

CausalLMModel

text-generation

hunyuan_v1_dense

HunYuanV1DenseCausalLMModel

text-generation

internlm2

InternLM2CausalLMModel

text-generation

llada

LLaDAModel

masked-diffusion

llama

CausalLMModel

text-generation

llama4_text

Llama4CausalLMModel

text-generation

mamba

MambaCausalLMModel

ssm-text-generation

mamba2

Mamba2CausalLMModel

ssm2-text-generation

minicpm

CausalLMModel

text-generation

minicpm3

CausalLMModel

text-generation

minimax

MiniMaxCausalLMModel

hybrid-text-generation

ministral

CausalLMModel

text-generation

ministral3

CausalLMModel

text-generation

mistral

CausalLMModel

text-generation

modernbert-decoder

ModernBertDecoderModel

text-generation

mpt

MPTCausalLMModel

text-generation

nanochat

NanoChatCausalLMModel

text-generation

nemotron

NemotronCausalLMModel

text-generation

olmo

OLMoCausalLMModel

text-generation

olmo2

OLMo2CausalLMModel

text-generation

olmo3

OLMo2CausalLMModel

text-generation

open-llama

CausalLMModel

text-generation

openelm

CausalLMModel

text-generation

persimmon

PersimmonCausalLMModel

text-generation

phi

PhiCausalLMModel

text-generation

phi3

Phi3CausalLMModel

text-generation

phi3small

Phi3SmallCausalLMModel

text-generation

qwen

QwenCausalLMModel

text-generation

qwen2

CausalLMModel

text-generation

qwen2_5_vl_text

Qwen25VLTextModel

text-generation

qwen2_vl_text

Qwen25VLTextModel

text-generation

qwen3

Qwen3CausalLMModel

text-generation

qwen3_5_text

Qwen35CausalLMModel

hybrid-text-generation

qwen3_5_vl_text

Qwen35VLTextModel

hybrid-text-generation

qwen3_tts_tokenizer_12hz

Qwen3TTSTokenizerV2Model

text-generation

qwen3_vl_text

Qwen3VLTextModel

text-generation

seed_oss

CausalLMModel

text-generation

shieldgemma2

Gemma2CausalLMModel

text-generation

smollm3

SmolLM3CausalLMModel

text-generation

solar_open

CausalLMModel

text-generation

stablelm

LayerNormCausalLMModel

text-generation

starcoder2

StarCoder2CausalLMModel

text-generation

yi

CausalLMModel

text-generation

zamba

CausalLMModel

text-generation

Mixture of Experts

Models that route tokens to a subset of expert MLPs.

model_type

Class

Task

arctic

MoECausalLMModel

text-generation

dbrx

MoECausalLMModel

text-generation

deepseek_v2

DeepSeekV3CausalLMModel

text-generation

deepseek_v2_moe

DeepSeekV3CausalLMModel

text-generation

deepseek_v3

DeepSeekV3CausalLMModel

text-generation

deepseek_v4

DeepSeekV4CausalLMModel

deepseek-v4

dots1

DeepSeekV3CausalLMModel

text-generation

ernie4_5_moe

Ernie45MoECausalLMModel

text-generation

flex_olmo

MoECausalLMModel

text-generation

glm4_moe

Glm4MoECausalLMModel

text-generation

glm4v_moe_text

Glm4MoECausalLMModel

text-generation

gpt_oss

GPTOSSCausalLMModel

text-generation

granitemoe

GraniteMoECausalLMModel

text-generation

granitemoeshared

GraniteMoECausalLMModel

text-generation

hunyuan_v1_moe

HunYuanMoEV1CausalLMModel

text-generation

jetmoe

JetMoeCausalLMModel

text-generation

longcat_flash

LongcatFlashCausalLMModel

text-generation

mixtral

MoECausalLMModel

text-generation

olmoe

MoECausalLMModel

text-generation

phimoe

Phi3MoECausalLMModel

text-generation

qwen2_moe

Qwen2MoECausalLMModel

text-generation

qwen3_5_moe

Qwen35MoECausalLMModel

hybrid-text-generation

qwen3_moe

MoECausalLMModel

text-generation

qwen3_next

Qwen3NextCausalLMModel

hybrid-text-generation

qwen3_omni_moe

MoECausalLMModel

text-generation

qwen3_vl_moe

MoECausalLMModel

text-generation

youtu

DeepSeekV3CausalLMModel

text-generation

Multimodal

Models that process images, audio, or other modalities alongside text.

model_type

Class

Task

aya_vision

LLaVAModel

vision-language

blip-2

Blip2Model

vision-language

chameleon

LLaVAModel

vision-language

cohere2_vision

LLaVAModel

vision-language

deepseek_vl

LLaVAModel

vision-language

deepseek_vl_hybrid

LLaVAModel

vision-language

deepseek_vl_v2

DeepSeekOCR2CausalLMModel

vision-language

florence2

LLaVAModel

vision-language

fuyu

LLaVAModel

vision-language

gemma3

Gemma3MultiModalModel

vision-language

gemma4

Gemma4Model

gemma4

gemma4_unified

Gemma4UnifiedModel

gemma4-unified

glm4v

LLaVAModel

vision-language

glm4v_moe

LLaVAModel

vision-language

got_ocr2

LLaVAModel

vision-language

hunyuan_vl_mot

HunYuanVLMoTModel

hunyuan-vl-mot

idefics2

LLaVAModel

vision-language

idefics3

LLaVAModel

vision-language

instructblip

LLaVAModel

vision-language

instructblipvideo

LLaVAModel

vision-language

internvl

InternVL2Model

vision-language

internvl2

InternVL2Model

vision-language

internvl_chat

InternVL2Model

vision-language

janus

LLaVAModel

vision-language

llava

LLaVAModel

vision-language

llava_next

LLaVAModel

vision-language

llava_next_video

LLaVAModel

vision-language

llava_onevision

LLaVAModel

vision-language

mistral3

LLaVAModel

vision-language

mllama

MllamaCausalLMModel

mllama-vision-language

molmo

LLaVAModel

vision-language

ovis2

LLaVAModel

vision-language

paligemma

LLaVAModel

vision-language

phi3_v

Phi3VModel

vision-language

phi4-siglip

Phi4SigLIPModel

vision-language

phi4_multimodal

Phi4MMMultiModalModel

phi4mm-multimodal

phi4mm

Phi4MMMultiModalModel

phi4mm-multimodal

pixtral

LLaVAModel

vision-language

qwen2_5_vl

Qwen25VLCausalLMModel

qwen-vl

qwen2_vl

Qwen2VLCausalLMModel

qwen-vl

qwen3_5

Qwen35VL3ModelCausalLMModel

hybrid-qwen-vl

qwen3_5_moe_vl

Qwen35MoEVL3ModelCausalLMModel

hybrid-qwen-vl

qwen3_5_vl

Qwen35VL3ModelCausalLMModel

hybrid-qwen-vl

qwen3_vl

Qwen3VL3ModelCausalLMModel

qwen-vl

qwen3_vl_single

Qwen3VLCausalLMModel

qwen3-vl-vision-language

smolvlm

LLaVAModel

vision-language

video_llava

LLaVAModel

vision-language

vipllava

LLaVAModel

vision-language

Speech-to-Text

Encoder-decoder models for speech recognition.

model_type

Class

Task

fastconformer_rnnt

EncDecRNNTModel

fastconformer-rnnt

fun_asr

FunASRForConditionalGeneration

fun-asr-speech-language

mms

Wav2Vec2ForCTCModel

ctc-asr

qwen3_asr

Qwen3ASRForConditionalGeneration

speech-language

qwen3_forced_aligner

Qwen3ASRForConditionalGeneration

speech-language

sensevoice_small

SenseVoiceSmallModel

audio-ctc

whisper

WhisperForConditionalGeneration

speech-to-text

Audio

Audio encoder models for feature extraction (Wav2Vec2, HuBERT, WavLM).

model_type

Class

Task

data2vec-audio

Wav2Vec2Model

audio-feature-extraction

hubert

Wav2Vec2Model

audio-feature-extraction

mctct

Wav2Vec2Model

audio-feature-extraction

musicgen

Wav2Vec2Model

audio-feature-extraction

qwen3_tts

Qwen3TTSForConditionalGeneration

tts

seamless_m4t

Wav2Vec2Model

audio-feature-extraction

seamless_m4t_v2

Wav2Vec2Model

audio-feature-extraction

sew

Wav2Vec2Model

audio-feature-extraction

sew-d

Wav2Vec2Model

audio-feature-extraction

speecht5

Wav2Vec2Model

audio-feature-extraction

unispeech

Wav2Vec2Model

audio-feature-extraction

unispeech-sat

Wav2Vec2Model

audio-feature-extraction

voxtral_encoder

Wav2Vec2Model

audio-feature-extraction

wav2vec2

Wav2Vec2Model

audio-feature-extraction

wav2vec2-bert

Wav2Vec2Model

audio-feature-extraction

wav2vec2-conformer

Wav2Vec2Model

audio-feature-extraction

wavlm

Wav2Vec2Model

audio-feature-extraction

encoder-only

Encoder-only models for embeddings and classification (BERT, RoBERTa).

model_type

Class

Task

distilbert

DistilBertModel

feature-extraction

encoder

Encoder-only transformer models (BERT family).

model_type

Class

Task

albert

BertModel

feature-extraction

bert

BertModel

feature-extraction

bros

BertModel

feature-extraction

camembert

BertModel

feature-extraction

data2vec-text

BertModel

feature-extraction

deberta

BertModel

feature-extraction

deberta-v2

BertModel

feature-extraction

electra

BertModel

feature-extraction

ernie

BertModel

feature-extraction

ernie_m

BertModel

feature-extraction

esm

BertModel

feature-extraction

flaubert

BertModel

feature-extraction

ibert

BertModel

feature-extraction

layoutlm

BertModel

feature-extraction

layoutlmv2

BertModel

feature-extraction

layoutlmv3

LayoutLMv3Model

feature-extraction

lilt

BertModel

feature-extraction

markuplm

BertModel

feature-extraction

mega

BertModel

feature-extraction

megatron-bert

BertModel

feature-extraction

mobilebert

BertModel

feature-extraction

modernbert

ModernBertModel

feature-extraction

mpnet

BertModel

feature-extraction

mra

BertModel

feature-extraction

nezha

BertModel

feature-extraction

nystromformer

BertModel

feature-extraction

qdqbert

BertModel

feature-extraction

rembert

BertModel

feature-extraction

roberta

BertModel

feature-extraction

roberta-prelayernorm

BertModel

feature-extraction

roc_bert

BertModel

feature-extraction

roformer

BertModel

feature-extraction

splinter

BertModel

feature-extraction

squeezebert

BertModel

feature-extraction

xlm-roberta

BertModel

feature-extraction

xlm-roberta-xl

BertModel

feature-extraction

xlnet

BertModel

feature-extraction

xmod

BertModel

feature-extraction

yoso

BertModel

feature-extraction

encoder-decoder

Encoder-decoder sequence-to-sequence models (BART, T5, mBART).

model_type

Class

Task

bart

BartForConditionalGeneration

seq2seq

bigbird_pegasus

BartForConditionalGeneration

seq2seq

blenderbot

BartForConditionalGeneration

seq2seq

blenderbot-small

BartForConditionalGeneration

seq2seq

fsmt

BartForConditionalGeneration

seq2seq

led

BartForConditionalGeneration

seq2seq

longt5

T5ForConditionalGeneration

seq2seq

m2m_100

BartForConditionalGeneration

seq2seq

marian

BartForConditionalGeneration

seq2seq

mbart

BartForConditionalGeneration

seq2seq

mt5

T5ForConditionalGeneration

seq2seq

mvp

BartForConditionalGeneration

seq2seq

nllb-moe

BartForConditionalGeneration

seq2seq

nllb_moe

BartForConditionalGeneration

seq2seq

pegasus

BartForConditionalGeneration

seq2seq

pegasus_x

BartForConditionalGeneration

seq2seq

plbart

BartForConditionalGeneration

seq2seq

prophetnet

BartForConditionalGeneration

seq2seq

switch_transformers

T5ForConditionalGeneration

seq2seq

t5

T5ForConditionalGeneration

seq2seq

trocr

TrOCRForConditionalGeneration

seq2seq

umt5

T5ForConditionalGeneration

seq2seq

xlm-prophetnet

BartForConditionalGeneration

seq2seq

causal-lm

Absolute positional embedding language models (GPT-2 style).

model_type

Class

Task

biogpt

GPT2CausalLMModel

text-generation

ctrl

CTRLCausalLMModel

text-generation

gpt-sw3

GPT2CausalLMModel

text-generation

gpt2

GPT2CausalLMModel

text-generation

gpt_bigcode

GPT2CausalLMModel

text-generation

gpt_neo

GPT2CausalLMModel

text-generation

imagegpt

GPT2CausalLMModel

text-generation

openai-gpt

GPT2CausalLMModel

text-generation

opt

OPTCausalLMModel

text-generation

xglm

GPT2CausalLMModel

text-generation

xlm

XLMCausalLMModel

text-generation

vision

Vision-only models for image classification and feature extraction.

model_type

Class

Task

beit

ViTModel

image-classification

blip

BlipVisionModel

image-classification

clip_vision_model

CLIPVisionModel

image-classification

cvt

ViTModel

image-classification

data2vec-vision

ViTModel

image-classification

deit

ViTModel

image-classification

dinov2

ViTModel

image-classification

dinov2_with_registers

ViTModel

image-classification

dinov3_vit

ViTModel

image-classification

hiera

ViTModel

image-classification

ijepa

ViTModel

image-classification

mobilevit

ViTModel

image-classification

mobilevitv2

ViTModel

image-classification

pvt

ViTModel

image-classification

pvt_v2

ViTModel

image-classification

sam2

Sam2VisionModel

image-classification

siglip

SigLIPVisionModel

image-classification

siglip2

SigLIPVisionModel

image-classification

siglip2_vision_model

SigLIPVisionModel

image-classification

siglip_vision_model

SigLIPVisionModel

image-classification

swin

ViTModel

image-classification

swin2sr

ViTModel

image-classification

swinv2

ViTModel

image-classification

vit

ViTModel

image-classification

vit_hybrid

ViTModel

image-classification

vit_mae

ViTModel

image-classification

vit_msn

ViTModel

image-classification

Depth Estimation

model_type

Class

Task

depth_anything

DepthAnythingForDepthEstimation

image-classification

Encoder

model_type

Class

Task

clip_text_model

CLIPTextModel

feature-extraction

Hybrid SSM+Attention

model_type

Class

Task

bamba

BambaCausalLMModel

hybrid-text-generation

granitemoehybrid

GraniteMoeHybridCausalLMModel

hybrid-text-generation

jamba

JambaCausalLMModel

hybrid-text-generation

nemotron_h

NemotronHCausalLMModel

hybrid-text-generation

zamba2

Zamba2CausalLMModel

hybrid-text-generation

Object Detection

model_type

Class

Task

yolos

YolosForObjectDetection

object-detection

Segmentation

model_type

Class

Task

segformer

SegformerForSemanticSegmentation

image-classification

Text

model_type

Class

Task

gemma4_text

Gemma4CausalLMModel

gemma4-text-generation

gemma4_unified_text

Gemma4CausalLMModel

gemma4-text-generation