[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-94692":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":9,"htmlUrl":10,"language":11,"languages":9,"totalLinesOfCode":9,"stars":12,"forks":13,"watchers":14,"openIssues":15,"contributorsCount":9,"subscribersCount":16,"size":16,"stars1d":16,"stars7d":16,"stars30d":17,"stars90d":16,"forks30d":16,"starsTrendScore":16,"compositeScore":18,"rankGlobal":9,"rankLanguage":9,"license":9,"archived":19,"fork":19,"defaultBranch":20,"hasWiki":19,"hasPages":19,"topics":21,"createdAt":9,"pushedAt":9,"updatedAt":41,"readmeContent":42,"aiSummary":43,"trendingCount":16,"starSnapshotCount":16,"syncStatus":17,"lastSyncTime":44,"discoverSource":45},94692,"Automodel","NVIDIA-NeMo\u002FAutomodel","NVIDIA-NeMo","🚀 Pytorch Distributed native training library for LLMs\u002FVLMs with OOTB Hugging Face support",null,"https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel","Python",819,251,10,155,0,2,48.4,false,"main",[22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40],"llm","vlm","finetuning","gemma3","llama","llama3","mistral","openai","qwen3","gpt-oss","qwen3-next","glm","kimi-k2","deepseek-v3-2","minimax-m2","gemma4","qwen3-6","deepseek-v4","agent","2026-08-24 04:01:22","\u003Cdiv align=\"center\">\n\n# 🚀 NeMo AutoModel\n\n\u003C\u002Fdiv>\n\n\u003Cdiv align=\"center\">\n\n\u003C!-- [![License](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FLicense-Apache%202.0-blue.svg)](https:\u002F\u002Fopensource.org\u002Flicenses\u002FApache-2.0) -->\n[![codecov](https:\u002F\u002Fcodecov.io\u002Fgithub\u002FNVIDIA-NeMo\u002FAutomodel\u002Fgraph\u002Fbadge.svg?token=4NMKZVOW2Z)](https:\u002F\u002Fcodecov.io\u002Fgithub\u002FNVIDIA-NeMo\u002FAutomodel)\n[![CICD NeMo](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Factions\u002Fworkflows\u002Fcicd-main.yml\u002Fbadge.svg)](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Factions\u002Fworkflows\u002Fcicd-main.yml)\n[![Python 3.10+](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fpython-3.10+-blue.svg)](https:\u002F\u002Fwww.python.org\u002Fdownloads\u002Frelease\u002Fpython-3100\u002F)\n[![Contributions Welcome](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fcontributions-welcome-brightgreen.svg)](CONTRIBUTING.md)\n[![GitHub Stars](https:\u002F\u002Fimg.shields.io\u002Fgithub\u002Fstars\u002FNVIDIA-NeMo\u002FAutomodel.svg?style=social&label=Star)](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fstargazers\u002F)\n\n\u003C!-- **Day-0 integration with Hugging Face models automating fine-tuning and pretraining with pytorch-native parallelism, custom-kernels and optimized recipes**\n**Pytorch DTensor‑native SPMD library for large‑scale training**-->\n\n[📖 Documentation](https:\u002F\u002Fdocs.nvidia.com\u002Fnemo\u002Fautomodel\u002Flatest\u002Findex.html) • [🔥 Ready-to-Use Recipes](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002F#supported-models) • [💡 Examples](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Ftree\u002Fmain\u002Fexamples) • [Model Coverage](https:\u002F\u002Fdocs.nvidia.com\u002Fnemo\u002Fautomodel\u002Flatest\u002Fmodel-coverage\u002Foverview.html) • [Performance](https:\u002F\u002Fdocs.nvidia.com\u002Fnemo\u002Fautomodel\u002Flatest\u002Fperformance-summary.html) • [🤝 Contributing](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002FCONTRIBUTING.md)\n\n\u003C\u002Fdiv>\n\n## 📣 News and Discussions\n- [08\u002F12\u002F2026][**Qwen3.8-2.4T-A95B**](https:\u002F\u002Fhuggingface.co\u002FQwen\u002FQwen3.8-2.4T-A95B) We now support full-parameter fine-tuning for `Qwen\u002FQwen3.8-2.4T-A95B` checkpoints. Check out the [HellaSwag EP32\u002FPP8 recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fqwen\u002Fqwen3_8_2_4t_a95b_hellaswag_ep32_pp8.yaml) and [model coverage page](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fdocs\u002Fmodel-coverage\u002Fllm\u002Fqwen\u002Fqwen3-8-2-4t-a95b.mdx).\n- [08\u002F12\u002F2026][**North Micro Vision**](https:\u002F\u002Fhuggingface.co\u002FCohereLabs\u002FNorth-Micro-Vision-Instruct) We now support LoRA fine-tuning for Cohere Labs' 2.4B-parameter native-resolution vision-language model. Check out the [RDR recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fcohere_micro_vision\u002Fnorth_micro_vision_rdr.yaml) and [model coverage page](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fdocs\u002Fmodel-coverage\u002Fvlm\u002Fcoherelabs\u002Fnorth-micro-vision.mdx).\n- [08\u002F10\u002F2026][**MuseGlimmer**](https:\u002F\u002Fhuggingface.co\u002Fmeta-models\u002FMuse-Glimmer-30B) We now support SFT\u002FLoRA the dense 30B MuseGlimmer vision-language model, including TP\u002FCP packed sequence recipes. Check out the [recipes](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Ftree\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fmuse_glimmer) and [model coverage page](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fdocs\u002Fmodel-coverage\u002Fvlm\u002Fmuse\u002Fmuse_glimmer.mdx).\n- [08\u002F08\u002F2026][**Nemotron 3.5 Lightning**](https:\u002F\u002Fhuggingface.co\u002Fnvidia\u002FNVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16) We now support LoRA fine-tuning for NVIDIA's 30B-A3B hybrid MoE model with multi-token prediction. Check out the [HellaSwag recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fnemotron\u002Fnemotron_nano_v3_5_lightning_hellaswag_peft.yaml) and [model coverage page](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fdocs\u002Fmodel-coverage\u002Fllm\u002Fnvidia\u002Fnemotron-h.mdx).\n- [07\u002F30\u002F2026][**Inkling-Small**](https:\u002F\u002Fhuggingface.co\u002Fthinkingmachines\u002FInkling-Small) We now support full-parameter fine-tuning for the 276B-parameter, 12B-active Inkling-Small model on 64 H100 GPUs. Check out the [MedPix EP64 recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Finkling\u002FInkling_small_medpix_ep64.yaml) and [model coverage page](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fdocs\u002Fmodel-coverage\u002Fvlm\u002Fthinkingmachines\u002Finkling.mdx).\n- [07\u002F29\u002F2026][**Kimi K3**](https:\u002F\u002Fhuggingface.co\u002Fmoonshotai\u002FKimi-K3) We now support full-parameter fine-tuning for Moonshot AI's 2.8T-parameter MoE model on NVIDIA GB200. Check out the [HellaSwag EP32\u002FPP8 recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fkimi\u002Fk3_hellaswag.yaml) on 256 GB200.\n- [07\u002F21\u002F2026][**Laguna S 2.1**](https:\u002F\u002Fhuggingface.co\u002Fpoolside\u002FLaguna-S-2.1) We now support finetuning Poolside's 118B-A8B Laguna S 2.1 MoE model. Check out our [HellaSwag EP16 recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Flaguna\u002Flaguna_s_2p1_hellaswag_ep16.yaml).\n- [07\u002F18\u002F2026][**Inkling VLM MoE**](https:\u002F\u002Fhuggingface.co\u002Fthinkingmachines\u002FInkling) We now support finetuning Inkling. Check out our [MedPix recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Finkling\u002Finkling_medpix.yaml).\n- [06\u002F27\u002F2026] \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Frefs\u002Fheads\u002Fmain\u002Fdocs\u002Fassets\u002Fspeculative-decoding-spark.svg\" width=\"136\" height=\"24\" alt=\"animated speculative decoding marker\" \u002F> [**Speculative decoding in NeMo AutoModel**](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fdocs\u002Fguides\u002Fspeculative\u002Feagle.mdx) Train target-aligned drafters end to end: EAGLE-1\u002F2\u002F3, P-EAGLE, DFlash, and DeepSeek's newly released [DSpark](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fspeculative\u002Fdspark\u002Fqwen3_0.6b_dspark.yaml). Huge thanks to [@Khazic](https:\u002F\u002Fgithub.com\u002Fkhazic) for the speculative decoding stack and [@kashif](https:\u002F\u002Fgithub.com\u002Fkashif) for DSpark support.\n- [06\u002F21\u002F2026][**GLM-5.2**](https:\u002F\u002Fhuggingface.co\u002Fzai-org\u002FGLM-5.2) We now support finetuning `zai-org\u002FGLM-5.2` with IndexShare DSA, optional TileLang sparse kernels, and long-context CP recipes. Check out our [32K long-context recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fglm\u002Fglm_5.2_tulu3_32k_tilelang_cp8.yaml) and [model coverage page](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fdocs\u002Fmodel-coverage\u002Fllm\u002Fthudm\u002Fglm5-moe-dsa.mdx).\n- [06\u002F15\u002F2026][**DiffusionGemma**](https:\u002F\u002Fhuggingface.co\u002Fgoogle\u002Fdiffusiongemma-26B-A4B-it). We now support fine-tuning the `google\u002Fdiffusiongemma-26B-A4B-it` model. Check out our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fdllm_sft\u002Fdiffusion_gemma_sft.yaml) and [guide](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fdocs\u002Fguides\u002Fdllm\u002Fdiffusiongemma.mdx).\n- [06\u002F12\u002F2026][**MiniMax M3**](https:\u002F\u002Fhuggingface.co\u002FMiniMaxAI\u002FMiniMax-M3) We now support finetuning MiniMax's MiniMax-M3. Check out our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fminimax_m3\u002Fminimax_m3_vl_sft_ep32pp4.yaml) and [guide](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fdocs\u002Fguides\u002Fvlm\u002Fminimax-m3.mdx).\n- [06\u002F04\u002F2026][**Nemotron-3 Ultra**](https:\u002F\u002Fhuggingface.co\u002Fnvidia\u002FNVIDIA-Nemotron-3-Ultra-550B-A55B-BF16) We now support finetuning NVIDIA's Nemotron 3 Ultra 550B A55B. Check out our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fnemotron\u002Fnemotron_ultra_v3_hellaswag_peft.yaml) and [guide](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fdocs\u002Fguides\u002Fllm\u002Fnemotron-3-ultra.md).\n- [06\u002F03\u002F2026][**Gemma 4 12B**](https:\u002F\u002Fhuggingface.co\u002Fgoogle\u002Fgemma-4-12B) We now support finetuning the dense `google\u002Fgemma-4-12B` model. Check out our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fgemma\u002Fgemma_4_12b_hellaswag.yaml).\n- [05\u002F27\u002F2026][**Step-3.7-Flash**](https:\u002F\u002Fhuggingface.co\u002Fstepfun-ai\u002FStep-3.7-Flash) We added model coverage for Stepfun AI's 198B-A13B MoE vision-language model, targeting image\u002Fvideo agentic developer workflows with a 256k context language backbone and 1.8B ViT vision tower. See the [model coverage page](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fdocs\u002Fmodel-coverage\u002Fvlm\u002Fstepfun-ai\u002Fstep-3-7.md).\n- [05\u002F19\u002F2026][**Ling 2.0**](https:\u002F\u002Fhuggingface.co\u002Fcollections\u002FinclusionAI\u002Fling-20) We now support finetuning the inclusionAI Ling 2.0 MoE family (`inclusionAI\u002FLing-mini-2.0`, `inclusionAI\u002FLing-flash-2.0`, and `inclusionAI\u002FLing-1T`), thanks to [@Hayden727](https:\u002F\u002Fgithub.com\u002FHayden727). Check out our [recipes](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Ftree\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fling).\n- [05\u002F17\u002F2026][**ERNIE 4.5**](https:\u002F\u002Fhuggingface.co\u002Fbaidu) and [**MiMo-V2-Flash**](https:\u002F\u002Fhuggingface.co\u002FXiaomiMiMo\u002FMiMo-V2-Flash) We now support finetuning `baidu\u002FERNIE-4.5-0.3B-PT`, `baidu\u002FERNIE-4.5-21B-A3B-PT`, and `XiaomiMiMo\u002FMiMo-V2-Flash`. Check out our ERNIE [dense recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fernie4_5\u002Fernie4_5_0p3b_hellaswag.yaml), ERNIE [MoE recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fernie4_5\u002Fernie4_5_21b_a3b_hellaswag.yaml), and MiMo [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fmimo_v2_flash\u002Fmimo_v2_flash_hellaswag.yaml).\n- [04\u002F29\u002F2026][**Mistral Medium 3.5**](https:\u002F\u002Fhuggingface.co\u002Fmistralai\u002FMistral-Medium-3.5-128B) We now support finetuning Mistral AI's 128B FP8-native VLM Mistral Medium 3.5. Check out our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fmistral3p5\u002Fmistral3p5_128b_medpix.yaml) and [guide](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fdocs\u002Fguides\u002Fvlm\u002Fmistral-medium-3-5.md).\n- [04\u002F28\u002F2026][**Nemotron-3-Nano-Omni**](https:\u002F\u002Fhuggingface.co\u002Fnvidia\u002FNemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16) We now support finetuning `nvidia\u002FNemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16`, NVIDIA's 30B-A3B omnimodal MoE (text · image · audio) with NemotronH hybrid Mamba+Attention backbone. Check out our [SFT recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fnemotron_omni\u002Fnemotron_omni_cord_v2.yaml), [LoRA recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fnemotron_omni\u002Fnemotron_omni_cord_v2_peft.yaml), and [guide](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fdocs\u002Fguides\u002Fvlm\u002Fnemotron-omni.md).\n- [04\u002F28\u002F2026][**Hy3-preview**](https:\u002F\u002Fhuggingface.co\u002Ftencent\u002FHy3-preview) We now support finetuning `tencent\u002FHy3-preview`, thanks to [@Khazic](https:\u002F\u002Fgithub.com\u002Fkhazic). Check out our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fhy_v3\u002Fhy3_preview_deepep.yaml).\n- [04\u002F25\u002F2026][**DeepSeek V4 Flash**](https:\u002F\u002Fhuggingface.co\u002Fdeepseek-ai\u002FDeepSeek-V4-Flash) We now support finetuning `deepseek-ai\u002FDeepSeek-V4-Flash`, thanks to [@Khazic](https:\u002F\u002Fgithub.com\u002Fkhazic). Check out our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fdeepseek_v4\u002Fdeepseek_v4_flash_hellaswag.yaml) and [guide](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fdocs\u002Fguides\u002Fllm\u002Fdsv4-flash.md).\n- [04\u002F22\u002F2026][**Qwen3.6-27B**](https:\u002F\u002Fhuggingface.co\u002FQwen\u002FQwen3.6-27B) We now support finetuning `Qwen\u002FQwen3.6-27B`. Check out our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fqwen3_5\u002Fqwen3_6_27b.yaml).\n- [04\u002F20\u002F2026][**Qwen-Image**](https:\u002F\u002Fhuggingface.co\u002FQwen\u002FQwen-Image) We now support finetuning `Qwen\u002FQwen-Image`, thanks to [@harshareddy832](https:\u002F\u002Fgithub.com\u002Fharshareddy832). Check out our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fdiffusion\u002Ffinetune\u002Fqwen_image_t2i_flow.yaml).\n- [04\u002F16\u002F2026][**Qwen3.6 MoE**](https:\u002F\u002Fhuggingface.co\u002FQwen\u002FQwen3.6-35B-A3B) We now support finetuning `Qwen\u002FQwen3.6-35B-A3B`. Check out our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fqwen3_5_moe\u002Fqwen3_6_35b.yaml).\n- [04\u002F16\u002F2026][**LLaVA-OneVision-1.5**](https:\u002F\u002Fhuggingface.co\u002Flmms-lab\u002FLLaVA-OneVision-1.5-4B-Instruct) We now support finetuning `lmms-lab\u002FLLaVA-OneVision-1.5-4B-Instruct`, thanks to [@vgauraha62](https:\u002F\u002Fgithub.com\u002Fvgauraha62). Check out our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fllava_onevision\u002Fllava_ov_1_5_4b_finetune.yaml).\n- [04\u002F12\u002F2026][**MiniMax-M2.7**](https:\u002F\u002Fhuggingface.co\u002FMiniMaxAI\u002FMiniMax-M2.7) We now support finetuning `MiniMaxAI\u002FMiniMax-M2.7`. Check out our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fminimax_m2\u002Fminimax_m2.7_hellaswag_pp.yaml).\n- [04\u002F07\u002F2026][**GLM-5.1**](https:\u002F\u002Fhuggingface.co\u002Fzai-org\u002FGLM-5.1) We now support finetuning `zai-org\u002FGLM-5.1`. GLM-5.1 is Zhipu AI's latest open-source MoE model featuring MLA + DeepSeek Sparse Attention. Check out our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fglm\u002Fglm_5.1_hellaswag_pp.yaml) and [discussion](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fdiscussions\u002F1719).\n- [04\u002F02\u002F2026][**Gemma 4**](https:\u002F\u002Fhuggingface.co\u002Fcollections\u002Fgoogle\u002Fgemma-4) We support fine-tuning for Gemma4 (2B, 4B, 31B, 26BA4B)! Check out our [recipes](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Ftree\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fgemma4).\n- [03\u002F30\u002F2026]**NeMo AutoModel** ships with **agent-friendly skills** in [skills\u002F](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Ftree\u002Fmain\u002Fskills) to help you with common development tasks (e.g., running a recipe, model onboarding, development) across the repo. We welcome PRs that improve existing skills or add new ones.\n- [03\u002F16\u002F2026][**Mistral Small 4**](https:\u002F\u002Fhuggingface.co\u002Fmistralai\u002FMistral-Small-4-119B-2603) We support fine-tuning for Mistral4 119B! Check out our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fmistral4\u002Fmistral4_medpix.yaml).\n- [03\u002F11\u002F2026][**Nemotron Super v3**](https:\u002F\u002Fhuggingface.co\u002Fnvidia\u002FNVIDIA-Nemotron-3-Super-120B-A12B-BF16) We support fine-tuning for `nvidia\u002FNVIDIA-Nemotron-3-Super-120B-A12B-BF16`. Check out our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fnemotron\u002Fnemotron_super_v3_hellaswag.yaml).\n- [03\u002F11\u002F2026][**GLM-5**](https:\u002F\u002Fhuggingface.co\u002Fzai-org\u002FGLM-5) We now support finetuning `zai-org\u002FGLM-5`. Check out our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fglm\u002Fglm_5_hellaswag_pp.yaml).\n- [03\u002F02\u002F2026][**Qwen3.5 small models**](https:\u002F\u002Fhuggingface.co\u002Fcollections\u002FQwen\u002Fqwen35) We support finetuning for Qwen3.5 small models 0.8B, 2B, 4B ([recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fqwen3_5\u002Fqwen3_5_4b.yaml)) and 9B ([recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fqwen3_5\u002Fqwen3_5_9b.yaml))\n- [02\u002F16\u002F2026][**Qwen3.5 MoE**](https:\u002F\u002Fhuggingface.co\u002Fcollections\u002FQwen\u002Fqwen35) We support finetuning for `Qwen\u002FQwen3.5-397B-A17B` ([recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fqwen3_5_moe\u002Fqwen3_5_moe_medpix.yaml)) and `Qwen\u002FQwen3.5-35B-A3B` ([recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fqwen3_5_moe\u002Fqwen3_5_35b.yaml))\n\n\u003Cdetails>\n\u003Csummary>Previous News\u003C\u002Fsummary>\n    \n- [02\u002F13\u002F2026] [**MiniMax-M2.5**](https:\u002F\u002Fhuggingface.co\u002FMiniMaxAI\u002FMiniMax-M2.5) We support finetuning for `MiniMaxAI\u002FMiniMax-M2.5`. Checkout our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fminimax_m2\u002Fminimax_m2.5_hellaswag_pp.yaml)\n- [02\u002F11\u002F2026] [**GLM-4.7-Flash**](https:\u002F\u002Fhuggingface.co\u002Fzai-org\u002FGLM-4.7-Flash) We now support finetuning GLM-4.7-Flash. Checkout our [packed sequence recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fglm\u002Fglm_4.7_flash_te_packed_sequence.yaml)\n- [02\u002F09\u002F2026] [**MiniMax-M2**](https:\u002F\u002Fhuggingface.co\u002FMiniMaxAI\u002FMiniMax-M2) We support finetuning for `MiniMaxAI\u002FMiniMax-M2`. Checkout our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002F5f63eb428bacf4146e9a5ae9949d58c5751df7b9\u002Fexamples\u002Fllm_finetune\u002Fminimax_m2\u002Fminimax_m2.1_hellaswag_pp.yaml)\n- [02\u002F06\u002F2026] [**Qwen3 VL 235B**](https:\u002F\u002Fhuggingface.co\u002FQwen\u002FQwen3-VL-235B-A22B-Instruct) We support finetuning for `Qwen\u002FQwen3-VL-235B-A22B-Instruct`. Checkout our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fqwen3\u002Fqwen3_vl_moe_235b.yaml)\n- [02\u002F06\u002F2026] [**GLM4.7**](https:\u002F\u002Fhuggingface.co\u002Fzai-org\u002FGLM-4.7) We now support finetuning GLM4.7. Checkout our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fglm\u002Fglm_4.7_te_deepep.yaml)\n- [02\u002F06\u002F2026] [**Step3.5-flash**](https:\u002F\u002Fhuggingface.co\u002Fstepfun-ai\u002FStep-3.5-Flash) is out! Finetune it with our [finetune recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fstepfun\u002Fstep_3.5_flash_hellaswag_pp.yaml)\n- [02\u002F05\u002F2026] [**DeepSeek-V3.2**](https:\u002F\u002Fhuggingface.co\u002Fdeepseek-ai\u002FDeepSeek-V3.2) is out! Checkout out [the finetune recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fdeepseek_v32\u002Fdeepseek_v32_hellaswag_pp.yaml)!\n- [02\u002F04\u002F2026] [**Kimi K2.5 VL**](https:\u002F\u002Fhuggingface.co\u002Fmoonshotai\u002FKimi-K2.5) is out! Finetune it with [NeMo AutoModel](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fdiscussions\u002F1161)\n- [01\u002F30\u002F2026] [**Kimi VL**](https:\u002F\u002Fhuggingface.co\u002Fmoonshotai\u002FKimi-VL-A3B-Instruct) We support fine-tuning for `moonshotai\u002FKimi-VL-A3B-Instruct`. Check out our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fkimi\u002Fkimi2vl_cordv2.yaml).\n- [01\u002F12\u002F2026] [**Nemotron Flash**](https:\u002F\u002Fhuggingface.co\u002Fnvidia\u002FNemotron-Flash-1B) We support fine-tuning for `nvidia\u002FNemotron-Flash-1B`. Check out our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fnemotron_flash\u002Fnemotron_flash_1b_squad.yaml).\n- [01\u002F12\u002F2026] [**Nemotron Parse**](https:\u002F\u002Fhuggingface.co\u002Fnvidia\u002FNVIDIA-Nemotron-Parse-v1.1) We support fine-tuning `nvidia\u002FNVIDIA-Nemotron-Parse-v1.1` ([recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fnemotron\u002Fnemotron_parse_v1_1.yaml), [tutorial](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Ftutorials\u002Fnemotron-parse\u002Ffinetune.ipynb) and [try on Brev](https:\u002F\u002Fbrev.nvidia.com\u002Flaunchable\u002Fdeploy\u002Fnow?launchableID=env-3C6LDKU2DfOvpVTFhjw3YQ4djPM)).\n- [01\u002F07\u002F2026] [**Devstral-Small**](https:\u002F\u002Fhuggingface.co\u002Fmistralai\u002FDevstral-Small-2-24B-Instruct-2512) We support fine-tuning for `mistralai\u002FDevstral-Small-2-24B-Instruct-2512`. Check out our [recipe](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fdevstral\u002Fdevstral2_small_2512_squad.yaml).\n- [12\u002F18\u002F2025] [**FunctionGemma**](https:\u002F\u002Fhuggingface.co\u002Fgoogle\u002Ffunctiongemma-270m-it) is out! Finetune it with [NeMo AutoModel](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fdocs\u002Fguides\u002Fllm\u002Ftoolcalling.md)!\n- [12\u002F15\u002F2025] [**NVIDIA-Nemotron-3-Nano-30B-A3B**](https:\u002F\u002Fhuggingface.co\u002Fnvidia\u002FNVIDIA-Nemotron-3-Nano-30B-A3B-FP8) is out! Finetune it with [NeMo AutoModel](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fdiscussions\u002F976)!\n- [11\u002F6\u002F2025] [Accelerating Large-Scale Mixture-of-Experts Training in PyTorch](https:\u002F\u002Fdeveloper.nvidia.com\u002Fblog\u002Faccelerating-large-scale-mixture-of-experts-training-in-pytorch\u002F)\n- [10\u002F6\u002F2025] [Enabling PyTorch Native Pipeline Parallelism for 🤗 Hugging Face Transformer Models](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fdiscussions\u002F589)\n- [9\u002F22\u002F2025] [Fine-tune Hugging Face Models Instantly with Day-0 Support with NVIDIA NeMo AutoModel](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fdiscussions\u002F477)\n- [9\u002F18\u002F2025] [🚀 NeMo Framework Now Supports Google Gemma 3n: Efficient Multimodal Fine-tuning Made Simple](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fdiscussions\u002F494)\n\n\u003C\u002Fdetails>\n\n## Overview\n\nNemo AutoModel is a Pytorch DTensor‑native SPMD open-source training library under [NVIDIA NeMo Framework](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo), designed to streamline and scale training and finetuning for LLMs, VLMs, diffusion models, and retrieval models. Designed for flexibility, reproducibility, and scale, NeMo AutoModel enables both small-scale experiments and massive multi-GPU, multi-node deployments for fast experimentation in research and production environments.\n\u003Cp align=\"center\">\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\">\u003Cpicture>\n    \u003Csource media=\"(prefers-color-scheme: light)\" srcset=\"https:\u002F\u002Fraw.githubusercontent.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Frefs\u002Fheads\u002Fmain\u002Fdocs\u002Fautomodel_diagram.png\">\n    \u003Cimg alt=\"AutoModel Logo\" src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Frefs\u002Fheads\u002Fmain\u002Fdocs\u002Fautomodel_diagram.png\">\n\u003C\u002Fpicture>\u003C\u002Fa>\n\u003C\u002Fp>\n\n\nWhat you can expect:\n\n- **Hackable** with a modular design that allows easy integration, customization, and quick research prototypes.\n- **Minimal ceremony**: YAML-driven recipes; override any field using CLI.\n- **High performance and flexibility** with custom kernels and DTensor support.\n- **Seamless integration** with Hugging Face for day-0 model support, ease of use, and wide range of supported models.\n- **Efficient resource management** using Kubernetes and Slurm, enabling scalable and flexible deployment across configurations.\n- **Documentation** with step-by-step guides and runnable examples.\n\n\u003C!-- Please refer to our design documents for more details on the architecture and design philosophy. -->\n\n\u003C!-- NeMo Framework is NVIDIA's GPU accelerated, end-to-end training framework for large language models (LLMs), multi-modal models and speech models. It enables seamless scaling of training (both pretraining and post-training) workloads from single GPU to thousand-node clusters for both 🤗Hugging Face\u002FPyTorch and Megatron models. It includes a suite of libraries and recipe collections to help users train models from end to end. The **AutoModel library (\"NeMo AutoModel\")** provides GPU-accelerated PyTorch training for 🤗Hugging Face models on **Day-0**. Users can start training and fine-tuning models instantly without conversion delays, scale effortlessly with PyTorch-native parallelisms, optimized custom kernels, and memory-efficient recipes-all while preserving the original checkpoint format for seamless use across the Hugging Face ecosystem. -->\n\n### Why PyTorch Distributed and SPMD\n\n- **One program, any scale**: The same training script runs on 1 GPU or 1000+ by changing the mesh.\n- **PyTorch Distributed native**: Partition model\u002Foptimizer states with `DeviceMesh` + placements (`Shard`, `Replicate`).\n- **SPMD first**: Parallelism is configuration. No model rewrites when scaling up or changing strategy.\n- **Decoupled concerns**: Model code stays pure PyTorch; parallel strategy lives in config.\n- **Composability**: Mix **tensor**, **sequence**, and **data** parallel by editing placements.\n- **Portability**: Fewer bespoke abstractions; easier to reason about failure modes and restarts.\n\u003C!-- - **Interoperability**: HF models\u002Ftokenizers\u002Foptimizers plug in directly; no format round‑trips. -->\n\n\u003C!-- ### Key Features -->\n\n\u003C!-- - **Mesh‑defined parallelism**: Compose tensor\u002Fsequence\u002Fpipeline\u002Fdata parallel by changing placements and sizes. -->\n\u003C!-- - **FSDP2 on DTensor**: Memory‑efficient sharding (HSDP included) for large scale training. -->\n\u003C!-- - **Pretraining, SFT & PEFT**: Day‑0 support for LLMs both regimes with shared configs\u002Futilities.\n- **Mixed precision**: BF16\u002FFP16\u002FFP8; sequence packing; optimized CUDA kernels. -->\n\u003C!-- - **Mesh‑aware DCP**: Sharded SafeTensors with merge\u002Freshard utilities; interoperable with HF. -->\n\u003C!-- - **Large-Scale Distributed Training**: Built-in FSDP2 and Megatron-FSDP for seamless multi-node scaling. -->\n\u003C!-- - **Vision-Language Model Ready**: Native support for VLMs (Qwen2-VL, Gemma-3-VL, etc). -->\n\u003C!-- - **Day-0 Hugging Face Support**: Instantly fine-tune any model from the Hugging Face Hub. -->\n\n\n## Table of Contents\n- [Feature Roadmap](#feature-roadmap)\n- [Getting Started](#getting-started)\n- [LLM](#llm-pre-training)\n  - [Pre-training](#llm-pre-training)\n  - [Supervised Fine-Tuning (SFT)](#llm-supervised-fine-tuning-sft)\n  - [Parameter-Efficient Fine-Tuning (PEFT)](#llm-parameter-efficient-fine-tuning-peft)\n- [VLM](#vlm-supervised-fine-tuning-sft)\n  - [Supervised Fine-Tuning (SFT)](#vlm-supervised-fine-tuning-sft)\n  - [Parameter-Efficient Fine-Tuning (PEFT)](#vlm-parameter-efficient-fine-tuning-peft)\n- [Supported Models](#supported-models)\n- [Performance](#performance)\n- [Interoperability](#-interoperability)\n- [Contributing](#-contributing)\n- [License](#-license)\n\n> TL;DR: SPMD turns “how to parallelize” into a *runtime layout choice*, not a code fork.\n\n## Feature Roadmap\n\n✅ _Available now ([v0.5.0](https:\u002F\u002Fpypi.org\u002Fproject\u002Fnemo-automodel\u002F0.5.0\u002F) \u002F [26.06 container](https:\u002F\u002Fcatalog.ngc.nvidia.com\u002Forgs\u002Fnvidia\u002Fcontainers\u002Fnemo-automodel\u002Ftags?version=26.06.00))_ | 🔜 _Planned for 26.08_\n\nHigh-throughput scalable training\n- ✅ **PyTorch DTensor-native SPMD training** Same training script can scale from 1 GPU to large multi-node jobs by changing the device mesh\u002Fconfig.\n- ✅ **Composable Parallelism** - PyTorch native FSDP2, HSDP, TP, CP, SP and PP for distributed training.\n- ✅ **Optimized kernels** - Uses NVIDIA-oriented kernel paths such as Transformer Engine, DeepEP, FlexAttn, TorchSDPA, fused attention, rotary embeddings, Triton, and optional kernel patches.\n- ✅ **MoE acceleration** - Includes MoE routing and DeepEP integration, plus expert-parallel configurations used in DeepSeek, Qwen MoE, GPT-OSS, and Nemotron MoE benchmarks.\n- ✅ **FP8 and mixed precision** - FP8 support with torchao and Transformer Engine.\n- ✅ **MXFP8 MoE training** - Transformer Engine and torchao MXFP8 grouped-expert training on GB200.\n- ✅ **Activation checkpointing** - Trades recomputation for lower activation memory, especially useful with FSDP and memory-efficient losses.\n- ✅ **Memory-efficient loss** - Linear-Cut \u002F fused linear cross entropy avoids materializing full logits for the loss, reducing output-layer memory pressure.\n- ✅ **Sequence packing** - Packs variable-length examples together to reduce padding compute and improve GPU utilization.\n- ✅ **FlashAttention packed-sequence support** - Packed masks can feed variable-length FlashAttention paths using per-document cu_seqlens.\n- ✅ **DCP** - Supports PyTorch DCP and SafeTensors, sharded and consolidated layouts, merge\u002Freshard utilities, and Hugging Face-compatible outputs.\n- ✅ **Async checkpointing** - Can write checkpoints in the background to reduce training stalls caused by I\u002FO.\n- ✅ **Dion and Muon optimizers** - Distributed optimizer integrations with typed recipe configuration.\n- ✅ **Environment Support** - SLURM, interactive, SkyPilot, and Kubernetes (via SkyPilot) launchers.\n\nSOTA algorithms\n- ✅ **Pre-training** - Support for model pre-training, including DeepSeekV3.\n- ✅ **Learning Algorithms** - SFT (Supervised Fine-Tuning), PEFT (LoRA, QLoRA), and QAT (Quantization-Aware Training).\n- ✅ **Knowledge distillation** - Support for knowledge distillation with LLMs.\n- ✅ **VLM knowledge distillation** - Chunked KD loss and a Qwen3.5-VL teacher-student recipe.\n- ✅ **Speculative decoding training** - Train target-aligned EAGLE-1\u002F2\u002F3, P-EAGLE, DFlash, and DSpark drafters for faster verified generation.\n\nModel Coverage and 🤗 Ecosystem compatibility\n- ✅ **Transformers v5 🤗** - Built on latest transformers with device-mesh driven parallelism.\n- ✅ **🤗 HuggingFace Integration** - Works with dense models (e.g., Qwen, Llama3, etc) and large MoEs (e.g., DSv3, DSv4).\n- ✅ **VLM** - Finetuning for VLMs (Qwen2.5\u002F3\u002F3.5\u002F3.6 VL, Gemma-3\u002F3n\u002F4 VL, Mistral 3.5\u002F4, LLaVA-OneVision-1.5, Kimi-VL, etc.).\n- ✅ **Omnimodal** - Finetuning for omnimodal MoE models (Nemotron-3-Nano-Omni, Qwen3-Omni).\n- ✅ **Diffusion** - Pretraining and LoRA finetuning for image\u002Fvideo diffusion models (Qwen-Image, FLUX, Wan2.1, Wan2.2-T2V-A14B, Hunyuan).\n- ✅ **dLLM** - Discrete diffusion LM finetuning (LLaDA, LLaDA2, Nemotron-Labs-Diffusion, DiffusionGemma).\n- ✅ **Retrieval** - Bi-encoder and cross-encoder training with in-batch negative sampling.\n- ✅ **Extended MoE support** - GPT-OSS, Kimi K3, Qwen3 \u002F Qwen3.5 \u002F Qwen3.6 MoE, Qwen-next, MiniMax-M2.x, GLM-4.7 \u002F GLM-5 \u002F GLM-5.1 \u002F GLM-5.2, DeepSeek V3.2 \u002F V4 \u002F V4-Flash, ERNIE 4.5, MiMo-V2-Flash, Ling 2.0, Hy3-preview.\n\nAgentic Development and UX\n- ✅ **Agent-friendly skills** - Curated [`skills\u002F`](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Ftree\u002Fmain\u002Fskills) for common dev tasks (recipe runs, model onboarding, CI).\n\nPlanned for 26.08\n- 🔜 **Unified Engine API and recipes** - Introduce a common engine and consolidate the LLM and VLM recipe paths.\n- 🔜 **Composable component configuration** - Complete the typed config and `.build()` refactor across data and remaining components.\n- 🔜 **Packed long-context training with CP** - Combine THD sequence packing with context parallelism, including DeepSeek V4 coverage.\n- 🔜 **Kernel and runtime upgrades** - Add partial CUDA graphs, evaluate FlashAttention 3\u002F4, and upgrade to DeepEP v2.\n- 🔜 **Multimodal retrieval** - Expand vision-language retrieval datasets and models and improve retriever performance.\n- 🔜 **Convergence and regression testing** - Add automated loss, evaluation, and performance gates to NeMo CI.\n- 🔜 **Transformers 5.12 alignment** - Upgrade the Transformers integration and restore affected model parity coverage.\n- 🔜 **Recipe consistency** - Align diffusion recipes with the shared AutoModel recipe and configuration patterns.\n\n\n## Getting Started\n\nWe recommend using **uv** for reproducible Python environments.\n\n```bash\n# Setup environment before running any recipes\nuv venv\n\n# Choose ONE:\nuv sync --frozen  # LLM recipes (default)\n# uv sync --frozen --extra vlm --extra vlm-media  # VLM recipes (Qwen\u002FMistral\u002FOmni need vlm-media for video\u002Fvision; fixes: ImportError: qwen_vl_utils is not installed)\n# uv sync --frozen --extra cuda  # Optional CUDA deps (e.g., Transformer Engine, Mamba SSM)\n# uv sync --frozen --extra cuda_source  # Optional bitsandbytes dependency\n# uv sync --frozen --extra all  # Most optional deps (includes `vlm` and `cuda`; NOTE: excludes media — add --extra media for video\u002Fimage decode)\n# uv sync --frozen --all-extras  # Everything (includes `fa`, `moe`, `media`, etc.)\n\n# One-off runs (examples):\n# uv run --extra vlm \u003Ccommand>\n# uv run --extra cuda \u003Ccommand>\n\nuv run python -c \"import nemo_automodel; print('NeMo AutoModel ready')\"\n```\n\n\n### Run a Recipe\nAll recipes are launched via the `automodel` CLI (or its short alias `am`). Each YAML config specifies the recipe class and all training parameters:\n```bash\n# LLM example: multi-GPU fine-tuning with FSDP2\nautomodel examples\u002Fllm_finetune\u002Fllama3_2\u002Fllama3_2_1b_hellaswag.yaml --nproc-per-node 8\n\n# VLM example: single-GPU fine-tuning (Gemma-3-VL) with LoRA\nautomodel examples\u002Fvlm_finetune\u002Fgemma3\u002Fgemma3_vl_4b_cord_v2_peft.yaml\n\n# Both commands also work with uv run:\nuv run automodel examples\u002Fllm_finetune\u002Fllama3_2\u002Fllama3_2_1b_hellaswag.yaml --nproc-per-node 8\n```\n\n> [!TIP]\n> **NeMo-Run submission:** The `cli` extra adds NeMo Run to the base package: `uv pip install \"nemo-automodel[cli]\"`. It is additive; the base package still installs its core training dependencies, including PyTorch.\n\n\n## LLM Pre-training\n### LLM Pre-training Single Node\nWe provide an example SFT experiment using the [FineWeb dataset](https:\u002F\u002Farxiv.org\u002Fabs\u002F2406.17557\u002F) with a nano-GPT model, ideal for quick experimentation on a single node.\n```sh\nautomodel examples\u002Fllm_pretrain\u002Fnanogpt_pretrain.yaml --nproc-per-node 8\n```\n\n\u003C!-- ### LLM Pre-training Multi Node -->\n\n## LLM Supervised Fine-Tuning (SFT)\nWe provide an example SFT experiment using the [SQuAD dataset](https:\u002F\u002Frajpurkar.github.io\u002FSQuAD-explorer\u002F).\n\n\u003C!-- Refer to `examples\u002Fllm_finetune\u002Fannotated.yaml` for a full list of parameters that can be overridden. -->\n\n### LLM SFT Single Node\n\nThe default SFT configuration is set to run on a single GPU. To start the experiment:\n\n```sh\nautomodel examples\u002Fllm_finetune\u002Fllama3_2\u002Fllama3_2_1b_squad.yaml\n```\n\nThis fine-tunes the `Llama3.2-1B` model on the SQuAD dataset using a single GPU.\n\nTo use multiple GPUs on a single node, add the `--nproc-per-node` argument:\n\n```sh\nautomodel examples\u002Fllm_finetune\u002Fllama3_2\u002Fllama3_2_1b_squad.yaml --nproc-per-node 8\n```\n\n### LLM SFT Multi Node\nTo launch on a SLURM cluster, copy the reference sbatch script and adapt it to your cluster:\n```sh\ncp slurm.sub my_cluster.sub\n# Edit my_cluster.sub — change CONFIG, #SBATCH directives, container, mounts, etc.\nsbatch my_cluster.sub\n```\n\nAll cluster-specific settings (nodes, GPUs, partition, container, mounts) live in your sbatch script.\nNeMo-Run (`nemo_run:`) sections are also supported -- see our\n[cluster guide](https:\u002F\u002Fdocs.nvidia.com\u002Fnemo\u002Fautomodel\u002Flatest\u002Flauncher\u002Fcluster.html) for details.\n\n## LLM Parameter-Efficient Fine-Tuning (PEFT)\n\nWe provide a PEFT example using the [HellaSwag dataset](https:\u002F\u002Frowanzellers.com\u002Fhellaswag\u002F).\n\n### LLM PEFT Single Node\n```bash\n# Memory-efficient SFT with LoRA\nautomodel examples\u002Fllm_finetune\u002Fllama3_2\u002Fllama3_2_1b_hellaswag_peft.yaml\n\n# Override any YAML parameter via the command line:\nautomodel examples\u002Fllm_finetune\u002Fllama3_2\u002Fllama3_2_1b_hellaswag_peft.yaml \\\n  --step_scheduler.local_batch_size 16\n```\n\n> [!NOTE]\n> Launching a multi-node PEFT example uses the same `sbatch slurm.sub` workflow as the SFT case above.\n\n\n## VLM Supervised Fine-Tuning (SFT)\n\nWe provide a VLM SFT example using Qwen2.5-VL for end-to-end fine-tuning on image-text data.\n\n### VLM SFT Single Node\n```bash\n# Qwen2.5-VL on 8 GPUs\nautomodel examples\u002Fvlm_finetune\u002Fqwen2_5\u002Fqwen2_5_vl_3b_rdr.yaml --nproc-per-node 8\n```\n\n## VLM Parameter-Efficient Fine-Tuning (PEFT)\n\nWe provide a VLM PEFT (LoRA) example for memory-efficient adaptation with Gemma3 VLM.\n\n### VLM PEFT Single Node\n```bash\n# Gemma-3-VL PEFT on 8 GPUs\nautomodel examples\u002Fvlm_finetune\u002Fgemma3\u002Fgemma3_vl_4b_medpix_peft.yaml --nproc-per-node 8\n```\n\n\n## Supported Models\nNeMo AutoModel provides native support for a wide range of models available on the Hugging Face Hub, enabling efficient fine-tuning for various domains. Below is a small sample of ready-to-use families (train as-is or swap any compatible 🤗 causal LM), you can specify nearly any LLM\u002FVLM model available on 🤗 hub:\n\n| Domain | Model Family | Model ID | Recipes |\n|--------|--------------|----------|---------|\n| **LLM** | **GPT-OSS** | [`GPT-OSS-20B`](https:\u002F\u002Fhuggingface.co\u002Fopenai\u002Fgpt-oss-20b) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fgpt_oss\u002Fgpt_oss_20b.yaml) |\n|  |  | [`GPT-OSS-120B`](https:\u002F\u002Fhuggingface.co\u002Fopenai\u002Fgpt-oss-120b) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fgpt_oss\u002Fgpt_oss_120b.yaml) |\n| **LLM** | **DeepSeek** | [`DeepSeek-V3`](https:\u002F\u002Fhuggingface.co\u002Fdeepseek-ai\u002FDeepSeek-V3) | [Pretrain](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_pretrain\u002Fdeepseekv3_pretrain.yaml) |\n| **LLM** | **Kimi K3** | [`moonshotai\u002FKimi-K3`](https:\u002F\u002Fhuggingface.co\u002Fmoonshotai\u002FKimi-K3) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fkimi\u002Fk3_hellaswag.yaml) |\n| **LLM** | **Moonlight** | [`Moonlight-16B-TE`](https:\u002F\u002Fhuggingface.co\u002Fmoonshotai\u002FMoonlight-16B-A3B) | [Pretrain](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_pretrain\u002Fmegatron_pretrain_moonlight_16b_te_slurm.yaml), [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fmoonlight\u002Fmoonlight_16b_te.yaml) |\n| **LLM** | **Nemotron-H** | [`nvidia\u002FNVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16`](https:\u002F\u002Fhuggingface.co\u002Fnvidia\u002FNVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16) | [LoRA](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fnemotron\u002Fnemotron_nano_v3_5_lightning_hellaswag_peft.yaml) |\n| **LLM** | **Ling 2.0** | [`inclusionAI\u002FLing-mini-2.0`](https:\u002F\u002Fhuggingface.co\u002FinclusionAI\u002FLing-mini-2.0) | [LoRA SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fling\u002Fling_mini_2_0_squad.yaml), [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fling\u002Fling_mini_2_0_sft.yaml) |\n|  |  | [`inclusionAI\u002FLing-flash-2.0`](https:\u002F\u002Fhuggingface.co\u002FinclusionAI\u002FLing-flash-2.0) | [LoRA SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fling\u002Fling_flash_2_0_lora.yaml), [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fling\u002Fling_flash_2_0_sft.yaml) |\n|  |  | [`inclusionAI\u002FLing-1T`](https:\u002F\u002Fhuggingface.co\u002FinclusionAI\u002FLing-1T) | [LoRA SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fling\u002Fling_1t_lora_pp.yaml), [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fling\u002Fling_1t_sft.yaml) |\n| **LLM** | **ERNIE 4.5** | [`baidu\u002FERNIE-4.5-0.3B-PT`](https:\u002F\u002Fhuggingface.co\u002Fbaidu\u002FERNIE-4.5-0.3B-PT) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fernie4_5\u002Fernie4_5_0p3b_hellaswag.yaml) |\n|  |  | [`baidu\u002FERNIE-4.5-21B-A3B-PT`](https:\u002F\u002Fhuggingface.co\u002Fbaidu\u002FERNIE-4.5-21B-A3B-PT) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fernie4_5\u002Fernie4_5_21b_a3b_hellaswag.yaml) |\n| **LLM** | **MiMo V2 Flash** | [`XiaomiMiMo\u002FMiMo-V2-Flash`](https:\u002F\u002Fhuggingface.co\u002FXiaomiMiMo\u002FMiMo-V2-Flash) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fmimo_v2_flash\u002Fmimo_v2_flash_hellaswag.yaml) |\n| **LLM** |  **LLaMA** | [`meta-llama\u002FLlama-3.2-1B`](https:\u002F\u002Fhuggingface.co\u002Fmeta-llama\u002FLlama-3.2-1B) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fllama3_2\u002Fllama3_2_1b_squad.yaml), [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fllama3_2\u002Fllama3_2_1b_hellaswag_peft.yaml) |\n| | | [`meta-llama\u002FLlama-3.2-3B-Instruct`](https:\u002F\u002Fhuggingface.co\u002Fmeta-llama\u002FLlama-3.2-3B-Instruct) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fllama3_2\u002Fllama_3_2_3b_instruct_squad.yaml), [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fllama3_2\u002Fllama_3_2_3b_instruct_squad_peft.yaml) |\n| | | [`meta-llama\u002FLlama-3.1-8B`](https:\u002F\u002Fhuggingface.co\u002Fmeta-llama\u002FLlama-3.1-8B) | [FP8](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fllama3_1\u002Fllama3_1_8b_hellaswag_fp8.yaml) |\n| | | [`meta-llama\u002FLlama-3.3-70B-Instruct`](https:\u002F\u002Fhuggingface.co\u002Fmeta-llama\u002FLlama-3.3-70B-Instruct) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fllama3_3\u002Fllama_3_3_70b_instruct_squad.yaml), [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fllama3_3\u002Fllama_3_3_70b_instruct_squad_peft.yaml) |\n| **LLM** | **Mistral** | [`mistralai\u002FMistral-7B-v0.1`](https:\u002F\u002Fhuggingface.co\u002Fmistralai\u002FMistral-7B-v0.1) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fmistral\u002Fmistral_7b_squad.yaml), [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fmistral\u002Fmistral_7b_squad_peft.yaml), [FP8](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fmistral\u002Fmistral_7b_hellaswag_fp8.yaml) |\n|  |  | [`mistralai\u002FMistral-Nemo-Base-2407`](https:\u002F\u002Fhuggingface.co\u002Fmistralai\u002FMistral-Nemo-Base-2407) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fmistral\u002Fmistral_nemo_2407_squad.yaml), [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fmistral\u002Fmistral_nemo_2407_squad_peft.yaml), [FP8](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fmistral\u002Fmistral_nemo_2407_hellaswag_fp8.yaml) |\n|  |  | [`mistralai\u002FMixtral-8x7B-Instruct-v0.1`](https:\u002F\u002Fhuggingface.co\u002Fmistralai\u002FMixtral-8x7B-Instruct-v0.1) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fmistral\u002Fmixtral-8x7b-v0-1_squad.yaml), [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fmistral\u002Fmixtral-8x7b-v0-1_squad_peft.yaml) |\n| **LLM** | **Qwen** | [`Qwen\u002FQwen2.5-7B`](https:\u002F\u002Fhuggingface.co\u002FQwen\u002FQwen2.5-7B) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fqwen\u002Fqwen2_5_7b_squad.yaml), [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fqwen\u002Fqwen2_5_7b_squad_peft.yaml), [FP8](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fqwen\u002Fqwen2_5_7b_hellaswag_fp8.yaml) |\n|  |  | [`Qwen\u002FQwen3-0.6B`](https:\u002F\u002Fhuggingface.co\u002FQwen\u002FQwen3-0.6B) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fqwen\u002Fqwen3_0p6b_hellaswag.yaml), [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fqwen\u002Fqwen3_0p6b_hellaswag_peft.yaml) |\n|  |  | [`Qwen\u002FQwen3.8-2.4T-A95B`](https:\u002F\u002Fhuggingface.co\u002FQwen\u002FQwen3.8-2.4T-A95B) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fqwen\u002Fqwen3_8_2_4t_a95b_hellaswag_ep32_pp8.yaml) |\n|  |  | [`Qwen\u002FQwQ-32B`](https:\u002F\u002Fhuggingface.co\u002FQwen\u002FQwQ-32B) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fqwen\u002Fqwq_32b_squad.yaml), [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fqwen\u002Fqwq_32b_squad_peft.yaml) |\n| **LLM** | **Gemma** | [`google\u002Fgemma-3-270m`](https:\u002F\u002Fhuggingface.co\u002Fgoogle\u002Fgemma-3-270m) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fgemma\u002Fgemma_3_270m_squad.yaml), [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fgemma\u002Fgemma_3_270m_squad_peft.yaml) |\n| | | [`google\u002Fgemma-2-9b-it`](https:\u002F\u002Fhuggingface.co\u002Fgoogle\u002Fgemma-2-9b-it) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fgemma\u002Fgemma_2_9b_it_squad.yaml), [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fgemma\u002Fgemma_2_9b_it_squad_peft.yaml), [FP8](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fgemma\u002Fgemma_2_9b_it_hellaswag_fp8.yaml) |\n| | | [`google\u002Fgemma-7b`](https:\u002F\u002Fhuggingface.co\u002Fgoogle\u002Fgemma-7b) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fgemma\u002Fgemma_7b_squad.yaml), [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fgemma\u002Fgemma_7b_squad_peft.yaml) |\n| **LLM** | **Phi** | [`microsoft\u002Fphi-2`](https:\u002F\u002Fhuggingface.co\u002Fmicrosoft\u002Fphi-2) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fphi\u002Fphi_2_squad.yaml), [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fphi\u002Fphi_2_squad_peft.yaml) |\n|  |  | [`microsoft\u002FPhi-3-mini-4k-instruct`](https:\u002F\u002Fhuggingface.co\u002Fmicrosoft\u002FPhi-3-mini-4k-instruct) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fphi\u002Fphi_3_mini_it_squad.yaml), [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fphi\u002Fphi_3_mini_it_squad_peft.yaml) |\n|  |  | [`microsoft\u002Fphi-4`](https:\u002F\u002Fhuggingface.co\u002Fmicrosoft\u002Fphi-4) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fphi\u002Fphi_4_squad.yaml), [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fphi\u002Fphi_4_squad_peft.yaml), [FP8](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fphi\u002Fphi_4_hellaswag_fp8.yaml) |\n| **LLM** | **Seed** | [`ByteDance-Seed\u002FSeed-Coder-8B-Instruct`](https:\u002F\u002Fhuggingface.co\u002FByteDance-Seed\u002FSeed-Coder-8B-Instruct) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fseed\u002Fseed_coder_8b_instruct_squad.yaml), [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fseed\u002Fseed_coder_8b_instruct_squad_peft.yaml), [FP8](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fseed\u002Fseed_coder_8b_instruct_hellaswag_fp8.yaml) |\n|  |  | [`ByteDance-Seed\u002FSeed-OSS-36B-Instruct`](https:\u002F\u002Fhuggingface.co\u002FByteDance-Seed\u002FSeed-OSS-36B-Instruct) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fseed\u002Fseed_oss_36B_hellaswag.yaml), [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fseed\u002Fseed_oss_36B_hellaswag_peft.yaml) |\n| **LLM** | **Baichuan** | [`baichuan-inc\u002FBaichuan2-7B-Chat`](https:\u002F\u002Fhuggingface.co\u002Fbaichuan-inc\u002FBaichuan2-7B-Chat) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fbaichuan\u002Fbaichuan_2_7b_squad.yaml), [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fbaichuan\u002Fbaichuan_2_7b_squad_peft.yaml), [FP8](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune\u002Fbaichuan\u002Fbaichuan_2_7b_mock_fp8.yaml) |\n| **VLM** | **Gemma** | [`google\u002Fgemma-3-4b-it`](https:\u002F\u002Fhuggingface.co\u002Fgoogle\u002Fgemma-3-4b-it) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fgemma3\u002Fgemma3_vl_4b_cord_v2.yaml), [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fgemma3\u002Fgemma3_vl_4b_cord_v2_peft.yaml) |\n|  |  | [`google\u002Fgemma-3n-e4b-it`](https:\u002F\u002Fhuggingface.co\u002Fgoogle\u002Fgemma-3n-e4b-it) | [SFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fgemma3n\u002Fgemma3n_vl_4b_medpix.yaml), [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fgemma3n\u002Fgemma3n_vl_4b_medpix_peft.yaml) |\n| **VLM** | **North Micro Vision** | [`CohereLabs\u002FNorth-Micro-Vision-Instruct`](https:\u002F\u002Fhuggingface.co\u002FCohereLabs\u002FNorth-Micro-Vision-Instruct) | [PEFT](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune\u002Fcohere_micro_vision\u002Fnorth_micro_vision_rdr.yaml) |\n\n> [!NOTE]\n> Check out more [LLM](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fllm_finetune) and [VLM](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002Fexamples\u002Fvlm_finetune) examples. Any causal LM on Hugging Face Hub can be used with the base recipe template, just overwrite `--model.pretrained_model_name_or_path \u003Cmodel-id>` in the CLI or in the YAML config.\n\n\n## Performance\n\nNeMo AutoModel achieves great training performance on NVIDIA GPUs. Below are highlights from our benchmark results:\n\n| Model | #GPUs | Seq Length | Model TFLOPs\u002Fsec\u002FGPU | Tokens\u002Fsec\u002FGPU | Kernel Optimizations |\n|-------|------:|-----------:|---------------------:|---------------:|----------------------|\n| DeepSeek V3 671B | 256 | 4096 | 250 | 1,002 | TE + DeepEP |\n| GPT-OSS 20B | 8 | 4096 | 279 | 13,058 | TE + DeepEP + FlexAttn |\n| Qwen3 MoE 30B | 8 | 4096 | 212 | 11,842 | TE + DeepEP |\n\nFor complete benchmark results including configuration details, see the [Performance Summary](docs\u002Fperformance-summary.md).\n\n\u003C!--\n## Mesh‑Aware Checkpointing\n\nAutoModel writes **Distributed Checkpoints (DCP)** with SafeTensors\nshards. Checkpoints carry partition metadata to:\n\n- **Merge** into a single HF‑compatible checkpoint for inference.\n- **Reshard** when loading onto a different mesh\u002Ftopology.\n\nYAML sketch:\n```yaml\ncheckpoint:\nenabled: true\ncheckpoint_dir: .\u002Fcheckpoints\nsave_consolidated: final\nmodel_save_format: safetensors\n``` -->\n\n## 🔌 Interoperability\n\n- **[NeMo RL](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FRL)**: Use AutoModel checkpoints directly as starting points for DPO\u002FRM\u002FGRPO pipelines.\n- **[Hugging Face](https:\u002F\u002Fgithub.com\u002Fhuggingface\u002Ftransformers)**: Train any LLM\u002FVLM from 🤗 without format conversion.\n- **[Megatron Bridge](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FMegatron-Bridge)**: Optional conversions to\u002Ffrom Megatron formats for specific workflows.\n\n\n## 🗂️ Project Structure\n\n```\nNeMo-Automodel\u002F\n├── cli\u002F                            # `automodel` \u002F `am` CLI entry-point\n│   └── app.py\n├── docker\u002F                         # Container build files\n├── docs\u002F                           # Documentation and guides\n├── examples\u002F\n│   ├── convergence\u002F                # Convergence test configs\n│   ├── diffusion\u002F                  # Diffusion pretrain\u002Ffinetune configs\n│   ├── dllm_sft\u002F                   # Discrete diffusion LM SFT configs\n│   ├── dllm_generate\u002F              # Discrete diffusion LM generation\n│   ├── llm_benchmark\u002F              # LLM benchmarking configs\n│   ├── llm_finetune\u002F               # LLM finetune YAML configs\n│   ├── llm_kd\u002F                     # LLM knowledge-distillation configs\n│   ├── llm_pretrain\u002F               # LLM pretrain configs\n│   ├── llm_seq_cls\u002F                # LLM sequence classification configs\n│   ├── retrieval\u002F                  # Bi-encoder \u002F cross-encoder configs\n│   ├── vlm_benchmark\u002F              # VLM benchmarking configs\n│   ├── vlm_finetune\u002F               # VLM finetune configs\n│   └── vlm_generate\u002F               # VLM generation configs\n├── nemo_automodel\u002F\n│   ├── _diffusers\u002F                 # HF Diffusers integration (NeMoAutoDiffusionPipeline)\n│   ├── _transformers\u002F              # HF Transformers integration\n│   ├── components\u002F                 # Core library\n│   │   ├── _peft\u002F                  # PEFT implementations (LoRA, QLoRA)\n│   │   ├── attention\u002F              # Attention implementations\n│   │   ├── checkpoint\u002F             # Distributed checkpointing\n│   │   ├── config\u002F\n│   │   ├── datasets\u002F               # LLM, VLM, diffusion, retrieval datasets\n│   │   ├── distributed\u002F            # FSDP2, Megatron FSDP, pipelining, CP, etc.\n│   │   ├── launcher\u002F               # Launcher backends (SLURM, NeMo-Run, SkyPilot)\n│   │   ├── loggers\u002F                # Loggers\n│   │   ├── loss\u002F                   # Optimized loss functions\n│   │   ├── models\u002F                 # User-defined model examples\n│   │   ├── moe\u002F                    # Optimized kernels for MoE models\n│   │   ├── optim\u002F                  # Optimizer\u002FLR scheduler components (incl. Dion)\n│   │   ├── quantization\u002F           # FP8, QAT, QLoRA\n│   │   ├── training\u002F               # Train utils\n│   │   └── utils\u002F                  # Misc utils\n│   ├── recipes\u002F\n│   │   ├── llm\u002F                    # Main LLM train loop\n│   │   ├── vlm\u002F                    # Main VLM train loop\n│   │   ├── diffusion\u002F              # Diffusion training loop\n│   │   ├── dllm\u002F                   # Discrete diffusion LM training loop\n│   │   └── retrieval\u002F              # Retrieval \u002F biencoder training loop\n│   └── shared\u002F\n├── tools\u002F                          # Developer tooling\n└── tests\u002F                          # Comprehensive test suite\n```\n\n\n## Citation\nIf you use NeMo AutoModel in your research, please cite it using the following BibTeX entry:\n```\n@misc{nemo-automodel,\ntitle = {NeMo AutoModel: DTensor-native SPMD library for scalable and efficient training},\nhowpublished = {\\url{https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel}},\nyear = {2025--2026},\nnote = {GitHub repository},\n}\n```\n\n## 🤝 Contributing\n\nWe welcome contributions! Please see our [Contributing Guide](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002FCONTRIBUTING.md) for details.\n\n\n## 📄 License\n\nNVIDIA NeMo AutoModel is licensed under the [Apache License 2.0](https:\u002F\u002Fgithub.com\u002FNVIDIA-NeMo\u002FAutomodel\u002Fblob\u002Fmain\u002FLICENSE).\n","NeMo AutoModel 是一个面向大语言模型（LLM）和多模态大模型（VLM）的 PyTorch 原生分布式训练库，支持开箱即用的 Hugging Face 模型集成。其核心功能包括零配置启动的全参数微调与预训练、基于 DTensor 的 SPMD 并行训练、定制化 CUDA 内核加速及预优化训练配方（如 LoRA、SFT、TP\u002FCP 序列打包）。技术特点涵盖对主流开源模型（如 Llama 3、Qwen3、Gemma、DeepSeek、MuseGlimmer 等）的深度适配与高性能扩展。适用于需要在多卡\u002F多节点环境下高效完成 LLM\u002FVLM 微调、指令对齐或领域适配的研究与工程场景。","2026-08-14 02:30:05","trending"]