# Supported Models mobius supports 298 registered model types. ## Text Generation Standard autoregressive language models (CausalLM). | `model_type` | Class | Task | |---|---|---| | {doc}`DFlashDraftModel` | `DFlashDraftModel` | `dflash-draft` | | {doc}`Eagle3DraftModel` | `Eagle3DraftModel` | `eagle3-draft` | | {doc}`Eagle3LlamaForCausalLM` | `Eagle3DraftModel` | `eagle3-draft` | | {doc}`Eagle3Speculator` | `Eagle3DraftModel` | `eagle3-draft` | | {doc}`Gemma4AssistantForCausalLM` | `Gemma4AssistantCausalLMModel` | `gemma4-assistant` | | {doc}`Gemma4UnifiedAssistantForCausalLM` | `Gemma4AssistantCausalLMModel` | `gemma4-assistant` | | {doc}`LlamaForCausalLMEagle3` | `Eagle3DraftModel` | `eagle3-draft` | | {doc}`Qwen35MtpModel` | `Qwen35MtpModel` | `qwen35-mtp` | | {doc}`apertus` | `ApertusCausalLMModel` | `text-generation` | | {doc}`arcee` | `ArceeCausalLMModel` | `text-generation` | | {doc}`baichuan` | `CausalLMModel` | `text-generation` | | {doc}`bloom` | `BloomCausalLMModel` | `text-generation` | | {doc}`chatglm` | `ChatGLMCausalLMModel` | `text-generation` | | {doc}`code_llama` | `CausalLMModel` | `text-generation` | | {doc}`codegen` | `CodeGenCausalLMModel` | `text-generation` | | {doc}`codegen2` | `CausalLMModel` | `text-generation` | | {doc}`cohere` | `CohereCausalLMModel` | `text-generation` | | {doc}`cohere2` | `CohereCausalLMModel` | `text-generation` | | {doc}`command_r` | `CausalLMModel` | `text-generation` | | {doc}`csm` | `CausalLMModel` | `text-generation` | | {doc}`diffllama` | `DiffLlamaCausalLMModel` | `text-generation` | | {doc}`doge` | `DogeCausalLMModel` | `text-generation` | | {doc}`ernie4_5` | `ErnieCausalLMModel` | `text-generation` | | {doc}`evolla` | `CausalLMModel` | `text-generation` | | {doc}`exaone` | `CausalLMModel` | `text-generation` | | {doc}`exaone4` | `ExaOne4CausalLMModel` | `text-generation` | | {doc}`falcon` | `FalconCausalLMModel` | `text-generation` | | {doc}`falcon_h1` | `FalconCausalLMModel` | `text-generation` | | {doc}`falcon_mamba` | `MambaCausalLMModel` | `ssm-text-generation` | | {doc}`gemma` | `GemmaCausalLMModel` | `text-generation` | | {doc}`gemma2` | `Gemma2CausalLMModel` | `text-generation` | | {doc}`gemma3_text` | `Gemma3CausalLMModel` | `text-generation` | | {doc}`gemma3n` | `Gemma3nCausalLMModel` | `text-generation` | | {doc}`gemma3n_text` | `Gemma3nCausalLMModel` | `text-generation` | | {doc}`gemma4_assistant` | `Gemma4AssistantCausalLMModel` | `gemma4-assistant` | | {doc}`gemma4_unified_assistant` | `Gemma4AssistantCausalLMModel` | `gemma4-assistant` | | {doc}`glm` | `GlmCausalLMModel` | `text-generation` | | {doc}`glm4` | `Glm4CausalLMModel` | `text-generation` | | {doc}`glm4v_text` | `Glm4CausalLMModel` | `text-generation` | | {doc}`gpt_neox` | `GPTNeoXCausalLMModel` | `text-generation` | | {doc}`gpt_neox_japanese` | `GPTNeoXJapaneseCausalLMModel` | `text-generation` | | {doc}`gptj` | `GPTJCausalLMModel` | `text-generation` | | {doc}`granite` | `GraniteCausalLMModel` | `text-generation` | | {doc}`helium` | `CausalLMModel` | `text-generation` | | {doc}`hunyuan_v1_dense` | `HunYuanV1DenseCausalLMModel` | `text-generation` | | {doc}`internlm2` | `InternLM2CausalLMModel` | `text-generation` | | {doc}`llada` | `LLaDAModel` | `masked-diffusion` | | {doc}`llama` | `CausalLMModel` | `text-generation` | | {doc}`llama4_text` | `Llama4CausalLMModel` | `text-generation` | | {doc}`mamba` | `MambaCausalLMModel` | `ssm-text-generation` | | {doc}`mamba2` | `Mamba2CausalLMModel` | `ssm2-text-generation` | | {doc}`minicpm` | `CausalLMModel` | `text-generation` | | {doc}`minicpm3` | `CausalLMModel` | `text-generation` | | {doc}`minimax` | `MiniMaxCausalLMModel` | `hybrid-text-generation` | | {doc}`ministral` | `CausalLMModel` | `text-generation` | | {doc}`ministral3` | `CausalLMModel` | `text-generation` | | {doc}`mistral` | `CausalLMModel` | `text-generation` | | {doc}`modernbert-decoder` | `ModernBertDecoderModel` | `text-generation` | | {doc}`mpt` | `MPTCausalLMModel` | `text-generation` | | {doc}`nanochat` | `NanoChatCausalLMModel` | `text-generation` | | {doc}`nemotron` | `NemotronCausalLMModel` | `text-generation` | | {doc}`olmo` | `OLMoCausalLMModel` | `text-generation` | | {doc}`olmo2` | `OLMo2CausalLMModel` | `text-generation` | | {doc}`olmo3` | `OLMo2CausalLMModel` | `text-generation` | | {doc}`open-llama` | `CausalLMModel` | `text-generation` | | {doc}`openelm` | `CausalLMModel` | `text-generation` | | {doc}`persimmon` | `PersimmonCausalLMModel` | `text-generation` | | {doc}`phi` | `PhiCausalLMModel` | `text-generation` | | {doc}`phi3` | `Phi3CausalLMModel` | `text-generation` | | {doc}`phi3small` | `Phi3SmallCausalLMModel` | `text-generation` | | {doc}`qwen` | `QwenCausalLMModel` | `text-generation` | | {doc}`qwen2` | `CausalLMModel` | `text-generation` | | {doc}`qwen2_5_vl_text` | `Qwen25VLTextModel` | `text-generation` | | {doc}`qwen2_vl_text` | `Qwen25VLTextModel` | `text-generation` | | {doc}`qwen3` | `Qwen3CausalLMModel` | `text-generation` | | {doc}`qwen3_5_text` | `Qwen35CausalLMModel` | `hybrid-text-generation` | | {doc}`qwen3_5_vl_text` | `Qwen35VLTextModel` | `hybrid-text-generation` | | {doc}`qwen3_tts_tokenizer_12hz` | `Qwen3TTSTokenizerV2Model` | `text-generation` | | {doc}`qwen3_vl_text` | `Qwen3VLTextModel` | `text-generation` | | {doc}`seed_oss` | `CausalLMModel` | `text-generation` | | {doc}`shieldgemma2` | `Gemma2CausalLMModel` | `text-generation` | | {doc}`smollm3` | `SmolLM3CausalLMModel` | `text-generation` | | {doc}`solar_open` | `CausalLMModel` | `text-generation` | | {doc}`stablelm` | `LayerNormCausalLMModel` | `text-generation` | | {doc}`starcoder2` | `StarCoder2CausalLMModel` | `text-generation` | | {doc}`yi` | `CausalLMModel` | `text-generation` | | {doc}`zamba` | `CausalLMModel` | `text-generation` | ## Mixture of Experts Models that route tokens to a subset of expert MLPs. | `model_type` | Class | Task | |---|---|---| | {doc}`arctic` | `MoECausalLMModel` | `text-generation` | | {doc}`dbrx` | `MoECausalLMModel` | `text-generation` | | {doc}`deepseek_v2` | `DeepSeekV3CausalLMModel` | `text-generation` | | {doc}`deepseek_v2_moe` | `DeepSeekV3CausalLMModel` | `text-generation` | | {doc}`deepseek_v3` | `DeepSeekV3CausalLMModel` | `text-generation` | | {doc}`deepseek_v4` | `DeepSeekV4CausalLMModel` | `deepseek-v4` | | {doc}`dots1` | `DeepSeekV3CausalLMModel` | `text-generation` | | {doc}`ernie4_5_moe` | `Ernie45MoECausalLMModel` | `text-generation` | | {doc}`flex_olmo` | `MoECausalLMModel` | `text-generation` | | {doc}`glm4_moe` | `Glm4MoECausalLMModel` | `text-generation` | | {doc}`glm4v_moe_text` | `Glm4MoECausalLMModel` | `text-generation` | | {doc}`gpt_oss` | `GPTOSSCausalLMModel` | `text-generation` | | {doc}`granitemoe` | `GraniteMoECausalLMModel` | `text-generation` | | {doc}`granitemoeshared` | `GraniteMoECausalLMModel` | `text-generation` | | {doc}`hunyuan_v1_moe` | `HunYuanMoEV1CausalLMModel` | `text-generation` | | {doc}`jetmoe` | `JetMoeCausalLMModel` | `text-generation` | | {doc}`longcat_flash` | `LongcatFlashCausalLMModel` | `text-generation` | | {doc}`mixtral` | `MoECausalLMModel` | `text-generation` | | {doc}`olmoe` | `MoECausalLMModel` | `text-generation` | | {doc}`phimoe` | `Phi3MoECausalLMModel` | `text-generation` | | {doc}`qwen2_moe` | `Qwen2MoECausalLMModel` | `text-generation` | | {doc}`qwen3_5_moe` | `Qwen35MoECausalLMModel` | `hybrid-text-generation` | | {doc}`qwen3_moe` | `MoECausalLMModel` | `text-generation` | | {doc}`qwen3_next` | `Qwen3NextCausalLMModel` | `hybrid-text-generation` | | {doc}`qwen3_omni_moe` | `MoECausalLMModel` | `text-generation` | | {doc}`qwen3_vl_moe` | `MoECausalLMModel` | `text-generation` | | {doc}`youtu` | `DeepSeekV3CausalLMModel` | `text-generation` | ## Multimodal Models that process images, audio, or other modalities alongside text. | `model_type` | Class | Task | |---|---|---| | {doc}`aya_vision` | `LLaVAModel` | `vision-language` | | {doc}`blip-2` | `Blip2Model` | `vision-language` | | {doc}`chameleon` | `LLaVAModel` | `vision-language` | | {doc}`cohere2_vision` | `LLaVAModel` | `vision-language` | | {doc}`deepseek_vl` | `LLaVAModel` | `vision-language` | | {doc}`deepseek_vl_hybrid` | `LLaVAModel` | `vision-language` | | {doc}`deepseek_vl_v2` | `DeepSeekOCR2CausalLMModel` | `vision-language` | | {doc}`florence2` | `LLaVAModel` | `vision-language` | | {doc}`fuyu` | `LLaVAModel` | `vision-language` | | {doc}`gemma3` | `Gemma3MultiModalModel` | `vision-language` | | {doc}`gemma4` | `Gemma4Model` | `gemma4` | | {doc}`gemma4_unified` | `Gemma4UnifiedModel` | `gemma4-unified` | | {doc}`glm4v` | `LLaVAModel` | `vision-language` | | {doc}`glm4v_moe` | `LLaVAModel` | `vision-language` | | {doc}`got_ocr2` | `LLaVAModel` | `vision-language` | | {doc}`hunyuan_vl_mot` | `HunYuanVLMoTModel` | `hunyuan-vl-mot` | | {doc}`idefics2` | `LLaVAModel` | `vision-language` | | {doc}`idefics3` | `LLaVAModel` | `vision-language` | | {doc}`instructblip` | `LLaVAModel` | `vision-language` | | {doc}`instructblipvideo` | `LLaVAModel` | `vision-language` | | {doc}`internvl` | `InternVL2Model` | `vision-language` | | {doc}`internvl2` | `InternVL2Model` | `vision-language` | | {doc}`internvl_chat` | `InternVL2Model` | `vision-language` | | {doc}`janus` | `LLaVAModel` | `vision-language` | | {doc}`llava` | `LLaVAModel` | `vision-language` | | {doc}`llava_next` | `LLaVAModel` | `vision-language` | | {doc}`llava_next_video` | `LLaVAModel` | `vision-language` | | {doc}`llava_onevision` | `LLaVAModel` | `vision-language` | | {doc}`mistral3` | `LLaVAModel` | `vision-language` | | {doc}`mllama` | `MllamaCausalLMModel` | `mllama-vision-language` | | {doc}`molmo` | `LLaVAModel` | `vision-language` | | {doc}`ovis2` | `LLaVAModel` | `vision-language` | | {doc}`paligemma` | `LLaVAModel` | `vision-language` | | {doc}`phi3_v` | `Phi3VModel` | `vision-language` | | {doc}`phi4-siglip` | `Phi4SigLIPModel` | `vision-language` | | {doc}`phi4_multimodal` | `Phi4MMMultiModalModel` | `phi4mm-multimodal` | | {doc}`phi4mm` | `Phi4MMMultiModalModel` | `phi4mm-multimodal` | | {doc}`pixtral` | `LLaVAModel` | `vision-language` | | {doc}`qwen2_5_vl` | `Qwen25VLCausalLMModel` | `qwen-vl` | | {doc}`qwen2_vl` | `Qwen2VLCausalLMModel` | `qwen-vl` | | {doc}`qwen3_5` | `Qwen35VL3ModelCausalLMModel` | `hybrid-qwen-vl` | | {doc}`qwen3_5_moe_vl` | `Qwen35MoEVL3ModelCausalLMModel` | `hybrid-qwen-vl` | | {doc}`qwen3_5_vl` | `Qwen35VL3ModelCausalLMModel` | `hybrid-qwen-vl` | | {doc}`qwen3_vl` | `Qwen3VL3ModelCausalLMModel` | `qwen-vl` | | {doc}`qwen3_vl_single` | `Qwen3VLCausalLMModel` | `qwen3-vl-vision-language` | | {doc}`smolvlm` | `LLaVAModel` | `vision-language` | | {doc}`video_llava` | `LLaVAModel` | `vision-language` | | {doc}`vipllava` | `LLaVAModel` | `vision-language` | ## Speech-to-Text Encoder-decoder models for speech recognition. | `model_type` | Class | Task | |---|---|---| | {doc}`fastconformer_rnnt` | `EncDecRNNTModel` | `fastconformer-rnnt` | | {doc}`fun_asr` | `FunASRForConditionalGeneration` | `fun-asr-speech-language` | | {doc}`mms` | `Wav2Vec2ForCTCModel` | `ctc-asr` | | {doc}`qwen3_asr` | `Qwen3ASRForConditionalGeneration` | `speech-language` | | {doc}`qwen3_forced_aligner` | `Qwen3ASRForConditionalGeneration` | `speech-language` | | {doc}`sensevoice_small` | `SenseVoiceSmallModel` | `audio-ctc` | | {doc}`whisper` | `WhisperForConditionalGeneration` | `speech-to-text` | ## Audio Audio encoder models for feature extraction (Wav2Vec2, HuBERT, WavLM). | `model_type` | Class | Task | |---|---|---| | {doc}`data2vec-audio` | `Wav2Vec2Model` | `audio-feature-extraction` | | {doc}`hubert` | `Wav2Vec2Model` | `audio-feature-extraction` | | {doc}`mctct` | `Wav2Vec2Model` | `audio-feature-extraction` | | {doc}`musicgen` | `Wav2Vec2Model` | `audio-feature-extraction` | | {doc}`qwen3_tts` | `Qwen3TTSForConditionalGeneration` | `tts` | | {doc}`seamless_m4t` | `Wav2Vec2Model` | `audio-feature-extraction` | | {doc}`seamless_m4t_v2` | `Wav2Vec2Model` | `audio-feature-extraction` | | {doc}`sew` | `Wav2Vec2Model` | `audio-feature-extraction` | | {doc}`sew-d` | `Wav2Vec2Model` | `audio-feature-extraction` | | {doc}`speecht5` | `Wav2Vec2Model` | `audio-feature-extraction` | | {doc}`unispeech` | `Wav2Vec2Model` | `audio-feature-extraction` | | {doc}`unispeech-sat` | `Wav2Vec2Model` | `audio-feature-extraction` | | {doc}`voxtral_encoder` | `Wav2Vec2Model` | `audio-feature-extraction` | | {doc}`wav2vec2` | `Wav2Vec2Model` | `audio-feature-extraction` | | {doc}`wav2vec2-bert` | `Wav2Vec2Model` | `audio-feature-extraction` | | {doc}`wav2vec2-conformer` | `Wav2Vec2Model` | `audio-feature-extraction` | | {doc}`wavlm` | `Wav2Vec2Model` | `audio-feature-extraction` | ## encoder-only Encoder-only models for embeddings and classification (BERT, RoBERTa). | `model_type` | Class | Task | |---|---|---| | {doc}`distilbert` | `DistilBertModel` | `feature-extraction` | ## encoder Encoder-only transformer models (BERT family). | `model_type` | Class | Task | |---|---|---| | {doc}`albert` | `BertModel` | `feature-extraction` | | {doc}`bert` | `BertModel` | `feature-extraction` | | {doc}`bros` | `BertModel` | `feature-extraction` | | {doc}`camembert` | `BertModel` | `feature-extraction` | | {doc}`data2vec-text` | `BertModel` | `feature-extraction` | | {doc}`deberta` | `BertModel` | `feature-extraction` | | {doc}`deberta-v2` | `BertModel` | `feature-extraction` | | {doc}`electra` | `BertModel` | `feature-extraction` | | {doc}`ernie` | `BertModel` | `feature-extraction` | | {doc}`ernie_m` | `BertModel` | `feature-extraction` | | {doc}`esm` | `BertModel` | `feature-extraction` | | {doc}`flaubert` | `BertModel` | `feature-extraction` | | {doc}`ibert` | `BertModel` | `feature-extraction` | | {doc}`layoutlm` | `BertModel` | `feature-extraction` | | {doc}`layoutlmv2` | `BertModel` | `feature-extraction` | | {doc}`layoutlmv3` | `LayoutLMv3Model` | `feature-extraction` | | {doc}`lilt` | `BertModel` | `feature-extraction` | | {doc}`markuplm` | `BertModel` | `feature-extraction` | | {doc}`mega` | `BertModel` | `feature-extraction` | | {doc}`megatron-bert` | `BertModel` | `feature-extraction` | | {doc}`mobilebert` | `BertModel` | `feature-extraction` | | {doc}`modernbert` | `ModernBertModel` | `feature-extraction` | | {doc}`mpnet` | `BertModel` | `feature-extraction` | | {doc}`mra` | `BertModel` | `feature-extraction` | | {doc}`nezha` | `BertModel` | `feature-extraction` | | {doc}`nystromformer` | `BertModel` | `feature-extraction` | | {doc}`qdqbert` | `BertModel` | `feature-extraction` | | {doc}`rembert` | `BertModel` | `feature-extraction` | | {doc}`roberta` | `BertModel` | `feature-extraction` | | {doc}`roberta-prelayernorm` | `BertModel` | `feature-extraction` | | {doc}`roc_bert` | `BertModel` | `feature-extraction` | | {doc}`roformer` | `BertModel` | `feature-extraction` | | {doc}`splinter` | `BertModel` | `feature-extraction` | | {doc}`squeezebert` | `BertModel` | `feature-extraction` | | {doc}`xlm-roberta` | `BertModel` | `feature-extraction` | | {doc}`xlm-roberta-xl` | `BertModel` | `feature-extraction` | | {doc}`xlnet` | `BertModel` | `feature-extraction` | | {doc}`xmod` | `BertModel` | `feature-extraction` | | {doc}`yoso` | `BertModel` | `feature-extraction` | ## encoder-decoder Encoder-decoder sequence-to-sequence models (BART, T5, mBART). | `model_type` | Class | Task | |---|---|---| | {doc}`bart` | `BartForConditionalGeneration` | `seq2seq` | | {doc}`bigbird_pegasus` | `BartForConditionalGeneration` | `seq2seq` | | {doc}`blenderbot` | `BartForConditionalGeneration` | `seq2seq` | | {doc}`blenderbot-small` | `BartForConditionalGeneration` | `seq2seq` | | {doc}`fsmt` | `BartForConditionalGeneration` | `seq2seq` | | {doc}`led` | `BartForConditionalGeneration` | `seq2seq` | | {doc}`longt5` | `T5ForConditionalGeneration` | `seq2seq` | | {doc}`m2m_100` | `BartForConditionalGeneration` | `seq2seq` | | {doc}`marian` | `BartForConditionalGeneration` | `seq2seq` | | {doc}`mbart` | `BartForConditionalGeneration` | `seq2seq` | | {doc}`mt5` | `T5ForConditionalGeneration` | `seq2seq` | | {doc}`mvp` | `BartForConditionalGeneration` | `seq2seq` | | {doc}`nllb-moe` | `BartForConditionalGeneration` | `seq2seq` | | {doc}`nllb_moe` | `BartForConditionalGeneration` | `seq2seq` | | {doc}`pegasus` | `BartForConditionalGeneration` | `seq2seq` | | {doc}`pegasus_x` | `BartForConditionalGeneration` | `seq2seq` | | {doc}`plbart` | `BartForConditionalGeneration` | `seq2seq` | | {doc}`prophetnet` | `BartForConditionalGeneration` | `seq2seq` | | {doc}`switch_transformers` | `T5ForConditionalGeneration` | `seq2seq` | | {doc}`t5` | `T5ForConditionalGeneration` | `seq2seq` | | {doc}`trocr` | `TrOCRForConditionalGeneration` | `seq2seq` | | {doc}`umt5` | `T5ForConditionalGeneration` | `seq2seq` | | {doc}`xlm-prophetnet` | `BartForConditionalGeneration` | `seq2seq` | ## causal-lm Absolute positional embedding language models (GPT-2 style). | `model_type` | Class | Task | |---|---|---| | {doc}`biogpt` | `GPT2CausalLMModel` | `text-generation` | | {doc}`ctrl` | `CTRLCausalLMModel` | `text-generation` | | {doc}`gpt-sw3` | `GPT2CausalLMModel` | `text-generation` | | {doc}`gpt2` | `GPT2CausalLMModel` | `text-generation` | | {doc}`gpt_bigcode` | `GPT2CausalLMModel` | `text-generation` | | {doc}`gpt_neo` | `GPT2CausalLMModel` | `text-generation` | | {doc}`imagegpt` | `GPT2CausalLMModel` | `text-generation` | | {doc}`openai-gpt` | `GPT2CausalLMModel` | `text-generation` | | {doc}`opt` | `OPTCausalLMModel` | `text-generation` | | {doc}`xglm` | `GPT2CausalLMModel` | `text-generation` | | {doc}`xlm` | `XLMCausalLMModel` | `text-generation` | ## vision Vision-only models for image classification and feature extraction. | `model_type` | Class | Task | |---|---|---| | {doc}`beit` | `ViTModel` | `image-classification` | | {doc}`blip` | `BlipVisionModel` | `image-classification` | | {doc}`clip_vision_model` | `CLIPVisionModel` | `image-classification` | | {doc}`cvt` | `ViTModel` | `image-classification` | | {doc}`data2vec-vision` | `ViTModel` | `image-classification` | | {doc}`deit` | `ViTModel` | `image-classification` | | {doc}`dinov2` | `ViTModel` | `image-classification` | | {doc}`dinov2_with_registers` | `ViTModel` | `image-classification` | | {doc}`dinov3_vit` | `ViTModel` | `image-classification` | | {doc}`hiera` | `ViTModel` | `image-classification` | | {doc}`ijepa` | `ViTModel` | `image-classification` | | {doc}`mobilevit` | `ViTModel` | `image-classification` | | {doc}`mobilevitv2` | `ViTModel` | `image-classification` | | {doc}`pvt` | `ViTModel` | `image-classification` | | {doc}`pvt_v2` | `ViTModel` | `image-classification` | | {doc}`sam2` | `Sam2VisionModel` | `image-classification` | | {doc}`siglip` | `SigLIPVisionModel` | `image-classification` | | {doc}`siglip2` | `SigLIPVisionModel` | `image-classification` | | {doc}`siglip2_vision_model` | `SigLIPVisionModel` | `image-classification` | | {doc}`siglip_vision_model` | `SigLIPVisionModel` | `image-classification` | | {doc}`swin` | `ViTModel` | `image-classification` | | {doc}`swin2sr` | `ViTModel` | `image-classification` | | {doc}`swinv2` | `ViTModel` | `image-classification` | | {doc}`vit` | `ViTModel` | `image-classification` | | {doc}`vit_hybrid` | `ViTModel` | `image-classification` | | {doc}`vit_mae` | `ViTModel` | `image-classification` | | {doc}`vit_msn` | `ViTModel` | `image-classification` | ## Depth Estimation | `model_type` | Class | Task | |---|---|---| | {doc}`depth_anything` | `DepthAnythingForDepthEstimation` | `image-classification` | ## Encoder | `model_type` | Class | Task | |---|---|---| | {doc}`clip_text_model` | `CLIPTextModel` | `feature-extraction` | ## Hybrid SSM+Attention | `model_type` | Class | Task | |---|---|---| | {doc}`bamba` | `BambaCausalLMModel` | `hybrid-text-generation` | | {doc}`granitemoehybrid` | `GraniteMoeHybridCausalLMModel` | `hybrid-text-generation` | | {doc}`jamba` | `JambaCausalLMModel` | `hybrid-text-generation` | | {doc}`nemotron_h` | `NemotronHCausalLMModel` | `hybrid-text-generation` | | {doc}`zamba2` | `Zamba2CausalLMModel` | `hybrid-text-generation` | ## Object Detection | `model_type` | Class | Task | |---|---|---| | {doc}`yolos` | `YolosForObjectDetection` | `object-detection` | ## Segmentation | `model_type` | Class | Task | |---|---|---| | {doc}`segformer` | `SegformerForSemanticSegmentation` | `image-classification` | ## Text | `model_type` | Class | Task | |---|---|---| | {doc}`gemma4_text` | `Gemma4CausalLMModel` | `gemma4-text-generation` | | {doc}`gemma4_unified_text` | `Gemma4CausalLMModel` | `gemma4-text-generation` | ```{toctree} :maxdepth: 1 :hidden: DFlashDraftModel Eagle3DraftModel Eagle3LlamaForCausalLM Eagle3Speculator Gemma4AssistantForCausalLM Gemma4UnifiedAssistantForCausalLM LlamaForCausalLMEagle3 Qwen35MtpModel albert apertus arcee arctic aya_vision baichuan bamba bart beit bert bigbird_pegasus biogpt blenderbot blenderbot-small blip blip-2 bloom bros camembert chameleon chatglm clip_text_model clip_vision_model code_llama codegen codegen2 cohere cohere2 cohere2_vision command_r csm ctrl cvt data2vec-audio data2vec-text data2vec-vision dbrx deberta deberta-v2 deepseek_v2 deepseek_v2_moe deepseek_v3 deepseek_v4 deepseek_vl deepseek_vl_hybrid deepseek_vl_v2 deit depth_anything diffllama dinov2 dinov2_with_registers dinov3_vit distilbert doge dots1 electra ernie ernie4_5 ernie4_5_moe ernie_m esm evolla exaone exaone4 falcon falcon_h1 falcon_mamba fastconformer_rnnt flaubert flex_olmo florence2 fsmt fun_asr fuyu gemma gemma2 gemma3 gemma3_text gemma3n gemma3n_text gemma4 gemma4_assistant gemma4_text gemma4_unified gemma4_unified_assistant gemma4_unified_text glm glm4 glm4_moe glm4v glm4v_moe glm4v_moe_text glm4v_text got_ocr2 gpt-sw3 gpt2 gpt_bigcode gpt_neo gpt_neox gpt_neox_japanese gpt_oss gptj granite granitemoe granitemoehybrid granitemoeshared helium hiera hubert hunyuan_v1_dense hunyuan_v1_moe hunyuan_vl_mot ibert idefics2 idefics3 ijepa imagegpt instructblip instructblipvideo internlm2 internvl internvl2 internvl_chat jamba janus jetmoe layoutlm layoutlmv2 layoutlmv3 led lilt llada llama llama4_text llava llava_next llava_next_video llava_onevision longcat_flash longt5 m2m_100 mamba mamba2 marian markuplm mbart mctct mega megatron-bert minicpm minicpm3 minimax ministral ministral3 mistral mistral3 mixtral mllama mms mobilebert mobilevit mobilevitv2 modernbert modernbert-decoder molmo mpnet mpt mra mt5 musicgen mvp nanochat nemotron nemotron_h nezha nllb-moe nllb_moe nystromformer olmo olmo2 olmo3 olmoe open-llama openai-gpt openelm opt ovis2 paligemma pegasus pegasus_x persimmon phi phi3 phi3_v phi3small phi4-siglip phi4_multimodal phi4mm phimoe pixtral plbart prophetnet pvt pvt_v2 qdqbert qwen qwen2 qwen2_5_vl qwen2_5_vl_text qwen2_moe qwen2_vl qwen2_vl_text qwen3 qwen3_5 qwen3_5_moe qwen3_5_moe_vl qwen3_5_text qwen3_5_vl qwen3_5_vl_text qwen3_asr qwen3_forced_aligner qwen3_moe qwen3_next qwen3_omni_moe qwen3_tts qwen3_tts_tokenizer_12hz qwen3_vl qwen3_vl_moe qwen3_vl_single qwen3_vl_text rembert roberta roberta-prelayernorm roc_bert roformer sam2 seamless_m4t seamless_m4t_v2 seed_oss segformer sensevoice_small sew sew-d shieldgemma2 siglip siglip2 siglip2_vision_model siglip_vision_model smollm3 smolvlm solar_open speecht5 splinter squeezebert stablelm starcoder2 swin swin2sr swinv2 switch_transformers t5 trocr umt5 unispeech unispeech-sat video_llava vipllava vit vit_hybrid vit_mae vit_msn voxtral_encoder wav2vec2 wav2vec2-bert wav2vec2-conformer wavlm whisper xglm xlm xlm-prophetnet xlm-roberta xlm-roberta-xl xlnet xmod yi yolos yoso youtu zamba zamba2 ```