[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-95885":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":9,"htmlUrl":10,"language":11,"languages":9,"totalLinesOfCode":9,"stars":12,"forks":13,"watchers":14,"openIssues":15,"contributorsCount":9,"subscribersCount":16,"size":16,"stars1d":16,"stars7d":16,"stars30d":17,"stars90d":16,"forks30d":16,"starsTrendScore":16,"compositeScore":18,"rankGlobal":9,"rankLanguage":9,"license":9,"archived":19,"fork":19,"defaultBranch":20,"hasWiki":19,"hasPages":19,"topics":9,"createdAt":9,"pushedAt":9,"updatedAt":21,"readmeContent":22,"aiSummary":23,"trendingCount":16,"starSnapshotCount":16,"syncStatus":24,"lastSyncTime":25,"discoverSource":26},95885,"miles","radixark\u002Fmiles","radixark","Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.",null,"https:\u002F\u002Fgithub.com\u002Fradixark\u002Fmiles","Python",2517,441,17,127,0,52,29.94,false,"main","2026-09-21 02:04:28","\u003Cdiv align=\"center\">\n\n\u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002Fradixark\u002Fmiles\u002Fmain\u002Fdocs\u002Fassets\u002Fimages\u002Fbrand\u002Fmiles_logo.png\" alt=\"Miles Logo\" width=\"340\">\n\n### **Enterprise-Grade Reinforcement Learning for Large-Scale Model Post-Training**\n\n[![Website](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fwebsite-miles.radixark.com-d55816)](https:\u002F\u002Fmiles.radixark.com\u002F)\n[![GitHub Repo](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fgithub-radixark%2Fmiles-black?logo=github)](https:\u002F\u002Fgithub.com\u002Fradixark\u002Fmiles)\n[![Docs](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fdocs-miles.radixark.com%2Fdocs-d55816)](https:\u002F\u002Fmiles.radixark.com\u002Fdocs)\n[![License](https:\u002F\u002Fimg.shields.io\u002Fgithub\u002Flicense\u002Fradixark\u002Fmiles)](LICENSE)\n[![Slack](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fslack-%23miles--rl-brightgreen.svg)](https:\u002F\u002Fslack.sglang.ai)\n\n| [**Website**](https:\u002F\u002Fmiles.radixark.com\u002F) | [**Documentation**](https:\u002F\u002Fmiles.radixark.com\u002Fdocs) | [**Quick Start**](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fgetting-started\u002Fquick-start) | [**Supported Models**](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fmodels) | [**Miles Diffusion**](https:\u002F\u002Fgithub.com\u002Fradixark\u002Fmiles_diffusion) | [**Blog**](https:\u002F\u002Fwww.lmsys.org\u002Fblog?filter=miles) | [**Slack**](https:\u002F\u002Fslack.sglang.ai) (`#miles-rl`) |\n\n\u003C\u002Fdiv>\n\n--------------------------------------------------------------------------------\n\n## News\n\n- [2026\u002F08] 🔥 Miles v0.1 is released! Read the blog post here: [Miles v0.1: Production-level Post-training](https:\u002F\u002Fwww.lmsys.org\u002Fblog\u002F2026-08-18-miles-v0-1).\n- [2026\u002F07] Towards Blackwell-Native 8-bit and 4-bit RL: End-to-End MXFP8 and NVFP4 RL in Miles ([blog](https:\u002F\u002Fwww.lmsys.org\u002Fblog\u002F2026-07-29-mxfp8-nvfp4-rl)).\n- [2026\u002F07] 🔥 SGLang and Miles add day-0 support for Kimi K3 ([blog](https:\u002F\u002Fwww.lmsys.org\u002Fblog\u002F2026-07-27-kimi-k3-day0-support)).\n- [2026\u002F07] On-policy distillation lands in Miles ([blog](https:\u002F\u002Fwww.lmsys.org\u002Fblog\u002F2026-07-18-opd-support-in-miles)).\n- [2026\u002F07] 🔥 SGLang and Miles add day-0 support for Inkling, a frontier multimodal model ([blog](https:\u002F\u002Fwww.lmsys.org\u002Fblog\u002F2026-07-15-inkling-day0-support)).\n- [2026\u002F07] DeepSeek-V4 Flash RL training comes to AMD Instinct MI355X with Miles ([blog](https:\u002F\u002Fwww.lmsys.org\u002Fblog\u002F2026-07-10-rocm-miles-dsv4)).\n- [2026\u002F06] SGLang and Miles add day-0 support for NVIDIA Nemotron 3 Ultra ([blog](https:\u002F\u002Fwww.lmsys.org\u002Fblog\u002F2026-06-04-nvidia-run-nemotron-3-ultra)).\n- [2026\u002F05] No token left behind: token-in-token-out in Miles ([blog](https:\u002F\u002Fwww.lmsys.org\u002Fblog\u002F2026-05-13-no-token-left-behind)).\n- [2026\u002F04] Updating 1 T parameters in seconds: P2P weight transfer in large-scale distributed RL ([blog](https:\u002F\u002Fwww.lmsys.org\u002Fblog\u002F2026-04-29-p2p-update)).\n- [2026\u002F04] 🔥 DeepSeek-V4 on day 0: from fast inference to verified RL with SGLang and Miles ([blog](https:\u002F\u002Fwww.lmsys.org\u002Fblog\u002F2026-04-25-deepseek-v4)).\n\n## About\n\nMiles is a high-performance, enterprise-ready reinforcement learning framework for\n**large-scale model post-training**. It pairs [SGLang](https:\u002F\u002Fgithub.com\u002Fsgl-project\u002Fsglang)\nfor high-throughput rollout with [Megatron-LM](https:\u002F\u002Fgithub.com\u002FNVIDIA\u002FMegatron-LM) for\nscalable training, and ships the precision, stability, and observability features an RL run\nneeds at trillion-parameter scale. A PyTorch FSDP2 backend is available for runs that would\nrather train the HuggingFace implementation as-is, though the recipes, the parallelism, and\nthe largest models all live on Megatron-LM. See\n[Training Backends](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fuser-guide\u002Ftraining-backend).\n\n> *\"A journey of a thousand miles begins with a single rollout.\"*\n\n### Performance\n\n- **Fully async RL.** Rollout and training workers are decoupled, with configurable on- and\n  off-policy schedules, a pipeline tuned for fewer bubbles, and customizable async rollout\n  and eval modes. See [Fully Async RL](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fuser-guide\u002Ffully-async).\n- **Fast agentic rollout.** Generation runs on [SGLang](https:\u002F\u002Fgithub.com\u002Fsgl-project\u002Fsglang)\n  behind a router that spreads requests across engines, preserves per-request metadata, and\n  health-checks the fleet. Tuned for multi-turn agentic workloads.\n- **Fast weight updates.** New weights reach the engines in-loop in seconds, even on a\n  trillion-parameter model such as Kimi-K2.6, with\n  [P2P RDMA](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fadvanced\u002Fp2p-weight-transfer) as the fast path\n  for disaggregated setups.\n- **Low-precision training.** [MXFP8 and NVFP4](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fadvanced\u002Flow-precision)\n  training with a numerically stable RL recipe that reduces precision-induced divergence.\n  FP8, [INT4 QAT](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fadvanced\u002Fint4-qat), BF16, and FP16 are also\n  supported.\n- **LoRA and multi-LoRA.** [Low-rank adapters](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fadvanced\u002Flora)\n  train frontier-scale models on a fraction of the GPUs, and the same adapters load straight\n  into SGLang for rollout.\n\n### Correctness and resilience\n\n- **Token-in-token-out (TITO).** Supported for\n  [every model and every black-box harness](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fuser-guide\u002Fagentic-rollout),\n  with no detokenize\u002Fretokenize round-trip between rollout and training.\n- **Rollout Routing Replay (R3).** Expert routing recorded during rollout is\n  [replayed in the trainer's forward pass](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fadvanced\u002Fmiles-router),\n  removing the MoE routing mismatch that destabilizes large runs, with compute and\n  communication overlapped to keep the cost down.\n- **Fault tolerance.** When an SGLang engine dies, Miles\n  [recovers it and resumes the run in place](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fadvanced\u002Ffault-tolerance):\n  no restart, no pause.\n\n### What Miles runs\n\n- **Day-0 model support.** DeepSeek-V4, Kimi-K3, GLM-5.2, Inkling, and Nemotron landed on\n  release day. Beyond day 0, nearly every frontier model runs on Miles, including Kimi-K2.6\n  and Qwen3.5. See [Models](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fmodels).\n- **Extensive hardware support.** NVIDIA GB300, GB200, B300, B200, H200, H100, and A100, and\n  AMD MI300X, MI325, MI350, and MI355X via ROCm. See\n  [Installation](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fgetting-started\u002Finstallation#hardware-requirements)\n  for per-GPU status and the container image for each.\n- **Wide recipe support.** GRPO, GSPO, PPO, and REINFORCE++ for RL, plus SFT and\n  [on-policy distillation](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fadvanced\u002Fon-policy-distillation).\n- **Agentic environments.** Train coding and computer-use agents through connectors for\n  Harbor, HUD, NeMo Gym, OpenEnv, Verifiers, and more, each plugging into the rollout\n  layer that fits it, with task sandboxes on AgentENV, Daytona, E2B, or Modal. See\n  [Agentic Environments](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fuser-guide\u002Fenvironments).\n- **Diffusion models.** Flow-GRPO, DiffusionNFT and SFT on an sglang-diffusion rollout\n  engine and an FSDP2 trainer, in\n  [Miles-diffusion](https:\u002F\u002Fgithub.com\u002Fradixark\u002Fmiles_diffusion).\n\n## Getting Started\n\n- [Install Miles](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fgetting-started\u002Finstallation)\n- [Quick Start](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fgetting-started\u002Fquick-start)\n- [Supported Models](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fmodels)\n- [Core Concepts](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fuser-guide\u002Fconcepts)\n- [Launch Script Walkthrough](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fuser-guide\u002Flaunch-script)\n- [Training Backends](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fuser-guide\u002Ftraining-backend)\n- [Contribution Guide](https:\u002F\u002Fmiles.radixark.com\u002Fdocs\u002Fdeveloper\u002Fcontributor-guide)\n\n## Acknowledgment\n\nMiles was forked from [slime](https:\u002F\u002Fgithub.com\u002FTHUDM\u002Fslime), and integrates\n[SGLang](https:\u002F\u002Fgithub.com\u002Fsgl-project\u002Fsglang),\n[Megatron-LM](https:\u002F\u002Fgithub.com\u002FNVIDIA\u002FMegatron-LM), and\n[torch_memory_saver](https:\u002F\u002Fgithub.com\u002Ffzyzcjy\u002Ftorch_memory_saver).\n\nMiles is shaped by the teams that build on it and support its development,\nfrom hardware and cloud to model labs, agent infrastructure, and academia:\n\n\u003Cdiv align=\"center\">\n\n\u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002Fradixark\u002Fmiles\u002Fmain\u002Fdocs\u002Fassets\u002Fimages\u002Facknowledgment.png\" alt=\"Organizations building on, contributing to, and collaborating with Miles\" width=\"900\">\n\n\u003C\u002Fdiv>\n\n## Citation\n\nIf Miles is useful in your research or your product, please cite it:\n\n```bibtex\n@misc{miles2026,\n  title        = {Miles: Enterprise-Grade Reinforcement Learning for Large-Scale Model Post-Training},\n  author       = {Miles Team},\n  year         = {2026},\n  howpublished = {\\url{https:\u002F\u002Fgithub.com\u002Fradixark\u002Fmiles}}\n}\n```\n","Miles 是一个面向企业的强化学习框架，专用于大语言模型（LLM）和多模态大模型（VLM）的后训练优化。它基于高性能 rollout 引擎 SGLang 和分布式训练库 Megatron-LM 构建，支持端到端 4\u002F8-bit 量化 RL、on-policy 蒸馏、token 级细粒度奖励建模及大规模分布式权重同步，具备低延迟推理与高吞吐训练能力。适用于需要在生产环境中对已预训练大模型进行安全对齐、指令微调、偏好优化或领域适配的企业级 AI 团队，尤其适合金融、客服、内容生成等需严格可控输出的场景。",2,"2026-09-05 02:30:09","trending"]