[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-94844":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":9,"htmlUrl":10,"language":11,"languages":9,"totalLinesOfCode":9,"stars":12,"forks":13,"watchers":14,"openIssues":14,"contributorsCount":9,"subscribersCount":14,"size":14,"stars1d":14,"stars7d":14,"stars30d":14,"stars90d":14,"forks30d":14,"starsTrendScore":14,"compositeScore":15,"rankGlobal":9,"rankLanguage":9,"license":9,"archived":16,"fork":16,"defaultBranch":17,"hasWiki":16,"hasPages":16,"topics":9,"createdAt":9,"pushedAt":9,"updatedAt":18,"readmeContent":19,"aiSummary":20,"trendingCount":14,"starSnapshotCount":14,"syncStatus":21,"lastSyncTime":9,"discoverSource":22},94844,"sglang-omni","sgl-project\u002Fsglang-omni","sgl-project","SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.",null,"https:\u002F\u002Fgithub.com\u002Fsgl-project\u002Fsglang-omni","Python",819,341,0,48.6,false,"main","2026-08-24 04:01:22","\u003Cdiv align=\"center\">\n\u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002Fsgl-project\u002Fsglang-omni\u002Fmain\u002Fdocs\u002F_static\u002Fimage\u002Fsgl-omni-logo.svg\" alt=\"logo\" width=\"400\">\u003C\u002Fimg>\n\n\u003Cp>\n\u003Ca href=\"https:\u002F\u002Fpypi.org\u002Fproject\u002Fsglang-omni\u002F\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fpypi\u002Fv\u002Fsglang-omni?style=for-the-badge&logo=pypi&logoColor=white&label=PyPI\" alt=\"PyPI\">\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fsgl-project\u002Fsglang-omni\u002Fstargazers\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fgithub\u002Fstars\u002Fsgl-project\u002Fsglang-omni?style=for-the-badge&logo=github&label=stars\" alt=\"GitHub stars\">\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fsgl-project\u002Fsglang-omni\u002Fblob\u002Fmain\u002FLICENSE\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fgithub\u002Flicense\u002Fsgl-project\u002Fsglang-omni?style=for-the-badge\" alt=\"license\">\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fsgl-project\u002Fsglang-omni\u002Fissues\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fgithub\u002Fissues-closed-raw\u002Fsgl-project\u002Fsglang-omni?style=for-the-badge&label=closed%20issues\" alt=\"closed issues\">\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fsgl-project\u002Fsglang-omni\u002Fissues\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fgithub\u002Fissues-raw\u002Fsgl-project\u002Fsglang-omni?style=for-the-badge&label=open%20issues\" alt=\"open issues\">\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fdeepwiki.com\u002Fsgl-project\u002Fsglang-omni\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FAsk-DeepWiki-087fca?style=for-the-badge\" alt=\"Ask DeepWiki\">\u003C\u002Fa>\n\u003C\u002Fp>\n\n\u003C\u002Fdiv>\n\n--------------------------------------------------------------------------------\n\n\u003Cp align=\"center\">\n\u003Ca href=\"https:\u002F\u002Flmsys.org\u002Fblog\u002F\">\u003Cb>Blog\u003C\u002Fb>\u003C\u002Fa> |\n\u003Ca href=\"https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002F\">\u003Cb>Documentation\u003C\u002Fb>\u003C\u002Fa> |\n\u003Ca href=\"#quick-start\">\u003Cb>Quick Start\u003C\u002Fb>\u003C\u002Fa> |\n\u003Ca href=\".\u002Fdocs\u002Fcookbook\u002F\">\u003Cb>Cookbook\u003C\u002Fb>\u003C\u002Fa> |\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fsgl-project\u002Fsglang\">\u003Cb>SGLang\u003C\u002Fb>\u003C\u002Fa> |\n\u003Ca href=\"https:\u002F\u002Fslack.sglang.io\">\u003Cb>Join Slack\u003C\u002Fb>\u003C\u002Fa>\n\u003C\u002Fp>\n\n\u003Cp align=\"center\">\n⭐ \u003Cb>\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fsgl-project\u002Fsglang-omni\u002Fstargazers\">Star SGLang-Omni\u003C\u002Fa> to help more builders discover open infrastructure for multimodal and speech serving!\u003C\u002Fb>\n\u003C\u002Fp>\n\n## News\n\n- [2026\u002F08] 🎵 Day-0 support for [MiniMax Music 3](https:\u002F\u002Fhuggingface.co\u002FMiniMaxAI\u002FMiniMax-Music3): lyrics + caption → 32 kHz stereo song on `\u002Fv1\u002Faudio\u002Fspeech`. \\[[Cookbook](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Fminimax_music3.html)\\]\n- [2026\u002F08] 🚀 SGLang-Omni **v0.1.2** is on [PyPI](https:\u002F\u002Fpypi.org\u002Fproject\u002Fsglang-omni\u002F). Install with `uv pip install \"sglang-omni==0.1.2\"`. \\[[Installation](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fget_started\u002Finstallation.html)\\] \\[[Release notes](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fget_started\u002Frelease_notes.html)\\]\n- [2026\u002F08] 🚀 TTS architecture refactor: shared pipeline state, engine construction, reference encoding, capability metadata, and vocoder scheduling. \\[[Roadmap](https:\u002F\u002Fgithub.com\u002Fsgl-project\u002Fsglang-omni\u002Fissues\u002F985)\\] \\[[Blog](https:\u002F\u002Fgithub.com\u002Fzhaochenyang20\u002FAwesome-ML-SYS-Tutorial\u002Fblob\u002Fmain\u002Fsglang\u002Fsglang-omni\u002Ftts-refactor.md)\\]\n- [2026\u002F06] 🔥 MOSS-TTS Local Transformer v1.5 on SGLang-Omni with native-streaming 48 kHz speech. \\[[Blog](https:\u002F\u002Flmsys.org\u002Fblog\u002F2026-06-17-moss-tts-local-v15\u002F)\\] \\[[Cookbook](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Fmoss_tts_local.html)\\]\n- [2026\u002F06] 🔥 Higgs Audio v3 TTS for real-time, controllable speech. \\[[Blog](https:\u002F\u002Flmsys.org\u002Fblog\u002F2026-06-04-higgs-audio-v3-tts\u002F)\\] \\[[Cookbook](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Fhiggs_tts.html)\\]\n\n## About\n\nSGLang-Omni is a multi-stage serving runtime for omni, speech, and TTS models. Its design target is multi-stage decoding: generation split across heterogeneous stages with different compute patterns, dependency structures, and resource needs. SGLang-Omni owns the pipeline topology, stage lifecycle, inter-stage transport, model-family integration layer, and OpenAI-compatible serving surface, while composing with [SGLang](https:\u002F\u002Fgithub.com\u002Fsgl-project\u002Fsglang) for high-performance autoregressive scheduling and model execution where applicable.\n\n- **Multi-stage runtime**: SGLang-Omni models generation as coordinated stages: preprocessing, encoders, autoregressive engines, talkers, decoders, vocoders, and aggregators.\n- **Stage-specialized scheduling**: Each stage runs behind a scheduler matched to its workload, from SGLang-backed autoregressive scheduling to lightweight preprocessing and streaming vocoder loops.\n- **Transport-aware execution**: A control plane coordinates requests while the relay data plane moves tensor payloads across shared-memory, NCCL, NIXL, and Mooncake backends.\n- **API surface**: OpenAI-compatible endpoints expose multimodal chat, speech generation, batch speech, streaming speech, uploaded voices, and transcription.\n\n## What SGLang-Omni Serves\n\n- **Omni chat and speech**: [Qwen3-Omni](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Fqwen3_omni.html), [Ming-Omni](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Fming_omni.html) — multimodal in, text\u002Faudio out.\n- **Music generation**: [MiniMax Music 3](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Fminimax_music3.html) — lyrics + caption → 32 kHz stereo song.\n- **Speech generation**: [Higgs Audio v3](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Fhiggs_tts.html), [MOSS-TTS](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Fmoss_tts.html), [MOSS-TTS Local](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Fmoss_tts_local.html), [Fish Speech S2-Pro](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Ffishaudio_s2_pro.html), [Qwen3-TTS](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Fqwen3_tts.html), [Voxtral TTS](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Fvoxtral_tts.html), [Ming-Omni-TTS](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Fming_tts.html), [dots.tts](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Fdots_tts.html), [ZONOS2](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Fzonos2.html) — `\u002Fv1\u002Faudio\u002Fspeech`, batch, streaming, uploaded voices.\n- **Audio transcription and diarization**: [Qwen3-ASR](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Fqwen3_asr.html), [Fun-ASR](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Ffun_asr.html), [ARK-ASR](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Farkasr.html), [MOSS-Transcribe-Diarize](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Fmoss_transcribe_diarize.html) via `\u002Fv1\u002Faudio\u002Ftranscriptions`. MOSS-TD supports speaker labels and timestamps (`response_format=verbose_json`).\n- **SGLang-Omni Router**: Multi-worker OpenAI-compatible front door — health, readiness, lifecycle, capability discovery. [Router guide](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fbasic_usage\u002Fomni_router.html).\n\n## Hardware Support\n\n| Backend | Status | Notes |\n|---------|--------|-------|\n| **NVIDIA CUDA** | ✅ Supported | Default target; full model coverage. |\n| **Intel GPU (XPU)** | 🧪 Experimental | Intel Arc GPUs via PyTorch XPU. **Qwen3-ASR, Qwen3-TTS, and Qwen3-Omni serve end-to-end** (Omni thinker via multi-XPU tensor parallelism). Install per [Intel XPU guide](.\u002Fdocs\u002Fget_started\u002Finstallation_xpu.md); the backend is auto-detected. |\n\nSee [Installation — Intel XPU](.\u002Fdocs\u002Fget_started\u002Finstallation_xpu.md).\n\nAdditional model guides, including experimental and research-oriented paths, are available in the [Cookbook](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002F).\n\n## Quick Start\n\n- [Installation](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fget_started\u002Finstallation.html)\n- [TTS usage](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fbasic_usage\u002Ftts.html)\n- [Qwen3-Omni usage](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fbasic_usage\u002Fqwen3_omni.html)\n- [Qwen3-ASR cookbook](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Fqwen3_asr.html)\n- [MOSS-Transcribe-Diarize cookbook](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fcookbook\u002Fmoss_transcribe_diarize.html)\n- [Omni router](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fbasic_usage\u002Fomni_router.html)\n- [Developer reference](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fdeveloper_reference\u002Fmain.html)\n\n## Community & Support\n\nSGLang-Omni welcomes contributors working on inference systems, kernels, scheduling, inter-stage communication, model runners and cache efficiency, model integration, benchmarking, production deployment. Join the [SGLang Slack](https:\u002F\u002Fslack.sglang.io) or read the [developer reference](https:\u002F\u002Fsgl-project.github.io\u002Fsglang-omni\u002Fdeveloper_reference\u002Fmain.html).\n\nOrganizations interested in supporting SGLang-Omni, TTS, or omni model serving can contact Chenyang Zhao at [zhaochenyang@lmsys.org](mailto:zhaochenyang@lmsys.org).\n\n## Acknowledgments\n\nSGLang-Omni builds on the SGLang ecosystem and on open model work from the TTS, speech, and omni-model communities. We thank the model teams, systems contributors, and partner organizations helping make open multimodal serving faster, more reliable, and easier to extend.\n","SGLang-Omni 是一个面向语音与多模态模型的高性能推理服务框架，专为 TTS（文本转语音）、ASR（自动语音识别）及通用语音\u002F多模态大模型提供低延迟、高吞吐的生产级部署支持。其核心特点包括统一服务接口、流式语音处理能力、共享状态管理、可扩展的编解码器调度机制，以及对 32–48 kHz 高保真音频的原生支持。适用于语音助手、实时字幕生成、AI 音乐合成、智能客服等需要高质量语音 I\u002FO 的工业级 AI 应用场景。",2,"trending"]