[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-95082":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":8,"htmlUrl":9,"language":10,"languages":8,"totalLinesOfCode":8,"stars":11,"forks":12,"watchers":13,"openIssues":14,"contributorsCount":8,"subscribersCount":15,"size":15,"stars1d":15,"stars7d":16,"stars30d":16,"stars90d":15,"forks30d":15,"starsTrendScore":16,"compositeScore":17,"rankGlobal":8,"rankLanguage":8,"license":8,"archived":18,"fork":18,"defaultBranch":19,"hasWiki":18,"hasPages":18,"topics":8,"createdAt":8,"pushedAt":8,"updatedAt":20,"readmeContent":21,"aiSummary":22,"trendingCount":15,"starSnapshotCount":15,"syncStatus":23,"lastSyncTime":24,"discoverSource":25},95082,"llm-scaler","intel\u002Fllm-scaler","intel",null,"https:\u002F\u002Fgithub.com\u002Fintel\u002Fllm-scaler","C++",490,71,36,98,0,1,40.17,false,"main","2026-08-24 04:01:23","# LLM Scaler\n\nLLM Scaler is an GenAI solution for text generation, image generation, video generation etc. running on Intel® Arc™ Pro B60 and B70 GPUs. LLM Scalar leverages standard frameworks such as vLLM, ComfyUI, SGLang Diffusion, Xinference etc and ensures the best performance for State-of-Art GenAI models running on Arc Pro B60\u002FB70 GPUs.\n\n---\n\n## Latest Update\n- 🔥[2026.08] We released `intel\u002Fllm-scaler-vllm:0.21.0-b3.1` for Muse-Glimmer-30B multi-modal support. \n- 🔥[2026.08] We released `intel\u002Fllm-scaler-omni:0.2.0-b1` to support ComfyUI 0.31 XPU stack, MiniMax H3 local video generation, Wan Animate 2 workflows, add optimizations for Wan 2.2 14B T2V Turbo, LTX-2, Z-Image\u002FLumina and Krea2 workflows, and support Quantized ComfyUI workflows (GGUF Q4_1 and Nunchaku W4A16)\n- 🔥[2026.08] We released `intel\u002Fllm-scaler-vllm:0.21.0-b3` to support Muse-Glimmer-30B, support DFlash for Muse-Glimmer-30B and Qwen3.6-27B, and improve TTFT for gemma-4-31B-it and gemma-4-26B-A4B-it. \n- 🔥[2026.08] We released `intel\u002Fllm-scaler-vllm:0.21.0-b2` to support Multi-token Prediction (MTP) and Lora Serving for Qwen3.6-27B, Qwen3.6-35B-A3B, gemma-4-31B-it and gemma-4-26B-A4B-it models, and support per-block quantization models Qwen3.6-27B-FP8 and Qwen3.6-35B-A3B-FP8. \n- [2026.07] We released `intel\u002Fllm-scaler-omni:0.1.0-b8` to support ComfyUI 0.27.0,more workflows and models.\n- [2026.07] We released `intel\u002Fllm-scaler-vllm:0.21.0-b1` to support gemma-4 (12B, 31B and 26B-A4B) and diffusiongemma (26B-A4B) models, and experimentally support XPU graph. \n- [2026.06] We released `intel\u002Fllm-scaler-vllm:0.14.0-b8.3.2` to fix Qwen3.5\u002F3.6-27B accuracy issues. \n- [2026.06] We released `intel\u002Fllm-scaler-vllm:0.14.0-b8.3.1` to enable FP8 KV Cache and fix bugs for Qwen3\u002FQwen3.5 models. \n- [2026.05] We released `intel\u002Fllm-scaler-vllm:0.14.0-b8.3` to improve performance for Qwen3.5\u002F3.6 series and Qwen3-Coder-Next, and enabled model streaming load to reduce peak memory. \n- [2026.05] We released `intel\u002Fllm-scaler-vllm:1.4` (or, `intel\u002Fllm-scaler-vllm:0.14.0-b8.2.1`) with new platform image and support Intel® Arc™ Pro B70 GPU. \n- [2026.05] We released `intel\u002Fllm-scaler-omni:0.1.0-b7` for more model workflows and performance improvments. \n- [2026.03] We released `intel\u002Fllm-scaler-vllm:0.14.0-b8.1` to support Qwen3.5-27B, Qwen3.5-35B-A3B and Qwen3.5-122B-A10B (FP8\u002FINT4 online quantization, GPTQ)\n- [2026.03] We released `intel\u002Fllm-scaler-omni:0.1.0-b6` for ComfyUI to support CacheDiT and torch.compile(), ComfyUI-GGUF, and more model workflows, and support FP8 for SGLang Diffusion.\n- [2026.03] We released `intel\u002Fllm-scaler-vllm:0.14.0-b8` for vLLM 0.14.0 and PyTorch 2.10 support, various new models support and performance improvement. \n- [2026.01] We released `intel\u002Fllm-scaler-vllm:1.3` (or, `intel\u002Fllm-scaler-vllm:0.11.1-b7`) for vLLM 0.11.1 and PyTorch 2.9 support, various new models support and performance improvement.\n- [2026.01] We released `intel\u002Fllm-scaler-omni:0.1.0-b5` for Python 3.12 and PyTorch 2.9 support, various ComfyUI workflows and more SGLang Diffusion support.\n- [2025.12] We released `intel\u002Fllm-scaler-vllm:1.2`, same image as `intel\u002Fllm-scaler-vllm:0.10.2-b6`. \n- [2025.12] We released `intel\u002Fllm-scaler-omni:0.1.0-b4` to support ComfyUI workflows for Z-Image-Turbo, Hunyuan-Video-1.5 T2V\u002FI2V with multi-XPU, and experimentially support SGLang Diffusion. \n- [2025.11] We released `intel\u002Fllm-scaler-vllm:0.10.2-b6` to support Qwen3-VL (Dense\u002FMoE), Qwen3-Omni, Qwen3-30B-A3B (MoE Int4), MinerU 2.5, ERNIE-4.5-vl etc. \n- [2025.11] We released `intel\u002Fllm-scaler-vllm:0.10.2-b5` to support gpt-oss models and released `intel\u002Fllm-scaler-omni:0.1.0-b3` to support more ComfyUI workflows, and Windows installation.\n- [2025.10] We released `intel\u002Fllm-scaler-omni:0.1.0-b2` to support more models with ComfyUI workflows and Xinference.\n- [2025.09] We released `intel\u002Fllm-scaler-vllm:0.10.0-b3` to support more models (MinerU, MiniCPM-v-4.5 etc), and released `intel\u002Fllm-scaler-omni:0.1.0-b1` to enable first omni GenAI models using ComfyUI and Xinference on Arc Pro B60 GPU.\n- [2025.08] We released `intel\u002Fllm-scaler-vllm:1.0`.\n\n\n\n## LLM Scaler vLLM\n\n`llm-scaler-vllm` supports running text generation models using vLLM, featuring: \n\n- ***CCL*** support (P2P or USM)\n- ***INT4*** and ***FP8*** quantized online serving, plus pre-quantized FP8 model support\n- ***Embedding*** and ***Reranker*** model support\n- ***Multi-Modal*** model support\n- ***Omni*** model support\n- ***Tensor Parallel***, ***Pipeline Parallel*** and ***Data Parallel***\n- Finding maximum Context Length\n- Multi-Modal WebUI\n- BPE-Qwen tokenizer\n\nPlease follow the instructions in the [Getting Started](vllm\u002FREADME.md\u002F#1-getting-started-and-usage) to use `llm-scaler-vllm`. \n\n### Supported Models\n\n\n| Model Name                                 | FP16 | Dynamic Online FP8 | Dynamic Online Int4 | MXFP4 | Notes                     |\n|--------------------------------------------|------|--------------------|----------------------|-------|---------------------------|\n| openai\u002Fgpt-oss-20b                         |      |                    |                      |   ✅   |                           |\n| openai\u002Fgpt-oss-120b                        |      |                    |                      |   ✅   |                           |\n| deepseek-ai\u002FDeepSeek-R1-Distill-Qwen-1.5B  |  ✅  |         ✅         |          ✅          |       |                           |\n| deepseek-ai\u002FDeepSeek-R1-Distill-Qwen-7B    |  ✅  |         ✅         |          ✅          |       |                           |\n| deepseek-ai\u002FDeepSeek-R1-Distill-Llama-8B   |  ✅  |         ✅         |          ✅          |       |                           |\n| deepseek-ai\u002FDeepSeek-R1-Distill-Qwen-14B   |  ✅  |         ✅         |          ✅          |       |                           |\n| deepseek-ai\u002FDeepSeek-R1-Distill-Qwen-32B   |  ✅  |         ✅         |          ✅          |       |                           |\n| deepseek-ai\u002FDeepSeek-R1-Distill-Llama-70B  |  ✅  |         ✅         |          ✅          |       |                           |\n| deepseek-ai\u002FDeepSeek-R1-0528-Qwen3-8B      |  ✅  |         ✅         |          ✅          |       |                           |\n| deepseek-ai\u002FDeepSeek-V2-Lite               |  ✅  |         ✅         |                      |       | export VLLM_MLA_DISABLE=1 |\n| deepseek-ai\u002Fdeepseek-coder-33b-instruct    |  ✅  |         ✅         |          ✅          |       |                           |\n| Qwen\u002FQwen3-8B                              |  ✅  |         ✅         |          ✅          |       |                           |\n| Qwen\u002FQwen3-14B                             |  ✅  |         ✅         |          ✅          |       |                           |\n| Qwen\u002FQwen3-32B                             |  ✅  |         ✅         |          ✅          |       |                           |\n| Qwen\u002FQwen3-30B-A3B                         |  ✅  |         ✅         |          ✅          |       |                           |\n| Qwen\u002FQwen3-235B-A22B                       |      |         ✅         |                      |       |                           |\n| Qwen\u002FQwen3-Coder-30B-A3B-Instruct          |  ✅  |         ✅         |          ✅          |       |                           |\n| Qwen\u002FQwen3-Coder-Next                      |  ✅  |         ✅         |                    |       |                           |\n| Qwen\u002FQwen3.5\u002F3.6-27B                       |  ✅  |         ✅         |          ✅          |       |                           |\n| Qwen\u002FQwen3.5\u002F3.6-35B-A3B                   |  ✅  |         ✅         |          ✅          |       |                           |\n| Qwen\u002FQwen3.6-27B-FP8                       |      |                    |                      |       | Pre-quantized offline FP8 model |\n| Qwen\u002FQwen3.6-35B-A3B-FP8                   |      |                    |                      |       | Pre-quantized offline FP8 model |\n| Qwen\u002FQwen3.5-122B-A10B                     |      |         ✅         |          ✅          |       |                           |\n| Qwen\u002FQwQ-32B                               |  ✅  |         ✅         |          ✅          |       |                           |\n| mistralai\u002FMinistral-8B-Instruct-2410       |  ✅  |         ✅         |          ✅          |       |                           |\n| mistralai\u002FMixtral-8x7B-Instruct-v0.1       |  ✅  |         ✅         |          ✅          |       |                           |\n| meta-llama\u002FLlama-3.1-8B                    |  ✅  |         ✅         |          ✅          |       |                           |\n| meta-llama\u002FLlama-3.1-70B                   |  ✅  |         ✅         |          ✅          |       |                           |\n| meta-models\u002FMuse-Glimmer-30B                   |     |         ✅         |                    |       |                           |\n| baichuan-inc\u002FBaichuan2-7B-Chat             |  ✅  |         ✅         |          ✅          |       | with chat_template        |\n| baichuan-inc\u002FBaichuan2-13B-Chat            |  ✅  |         ✅         |          ✅          |       | with chat_template        |\n| THUDM\u002FCodeGeex4-All-9B                     |  ✅  |         ✅         |          ✅          |       | with chat_template        |\n| zai-org\u002FGLM-4-9B-0414                      |      |         ✅        |                      |       | use bfloat16 |\n| zai-org\u002FGLM-4-32B-0414                     |      |         ✅        |                      |       | use bfloat16 |\n| zai-org\u002FGLM-4.5-Air                        |  ✅  |         ✅         |                      |       |                           |\n| zai-org\u002FGLM-4.7-Flash                      |  ✅  |         ✅         |                      |       |                           |\n| ByteDance-Seed\u002FSeed-OSS-36B-Instruct       |  ✅  |         ✅         |          ✅          |       |                           |\n| miromind-ai\u002FMiroThinker-v1.5-30B           |  ✅  |         ✅         |          ✅          |       |                           |\n| tencent\u002FHunyuan-0.5B-Instruct              |  ✅  |         ✅         |          ✅          |       |  follow the guide in [here](.\u002Fvllm\u002FREADME.md#31-how-to-use-hunyuan-7b-instruct)   |\n| tencent\u002FHunyuan-7B-Instruct                |  ✅  |         ✅         |          ✅          |       |  follow the guide in [here](.\u002Fvllm\u002FREADME.md#31-how-to-use-hunyuan-7b-instruct)   |\n| Qwen\u002FQwen2-VL-7B-Instruct                  |  ✅  |         ✅         |          ✅          |       |                           |\n| Qwen\u002FQwen2.5-VL-7B-Instruct                |  ✅  |         ✅         |          ✅          |       |                           |\n| Qwen\u002FQwen2.5-VL-32B-Instruct               |  ✅  |         ✅         |          ✅          |       |                           |\n| Qwen\u002FQwen2.5-VL-72B-Instruct               |  ✅  |         ✅         |          ✅          |       |                           |\n| Qwen\u002FQwen3-VL-4B-Instruct                  |  ✅  |         ✅         |          ✅          |       |                           |\n| Qwen\u002FQwen3-VL-8B-Instruct                  |  ✅  |         ✅         |          ✅          |       |                           |\n| Qwen\u002FQwen3-VL-30B-A3B-Instruct             |  ✅  |         ✅         |          ✅          |       |                           |\n| openbmb\u002FMiniCPM-V-2_6                      |  ✅  |         ✅         |          ✅          |       |                           |\n| openbmb\u002FMiniCPM-V-4                        |  ✅  |         ✅         |          ✅          |       |                           |\n| openbmb\u002FMiniCPM-V-4_5                      |  ✅  |         ✅         |          ✅          |       |                           |\n| OpenGVLab\u002FInternVL2-8B                     |  ✅  |         ✅         |          ✅          |       |                           |\n| OpenGVLab\u002FInternVL3-8B                     |  ✅  |         ✅         |          ✅          |       |                           |\n| OpenGVLab\u002FInternVL3_5-8B                   |  ✅  |         ✅         |          ✅          |       |                           |\n| OpenGVLab\u002FInternVL3_5-30B-A3B              |  ✅  |         ✅         |          ✅          |       |                           |\n| rednote-hilab\u002Fdots.ocr                     |  ✅  |         ✅         |          ✅          |       |                           |\n| ByteDance-Seed\u002FUI-TARS-7B-DPO              |  ✅  |         ✅         |          ✅          |       |                           |\n| google\u002Fgemma-3-12b-it                      |      |         ✅         |                      |       |  use bfloat16  |\n| google\u002Fgemma-3-27b-it                      |      |         ✅         |                      |       |  use bfloat16  |\n| google\u002Fgemma-4-12B-it                      |      |         ✅         |          ✅         |       |see [Reference Commands](vllm\u002FREADME.md\u002F#33-reference-commands-for-running-gemma-4-models-and-diffusiongemma)                            |\n| google\u002Fgemma-4-31B-it                      |      |         ✅         |          ✅         |       |                            |\n| google\u002Fgemma-4-26B-A4B-it                  |      |         ✅         |          ✅         |       |                            |\n| google\u002Fdiffusiongemma-26B-A4B-it           |      |         ✅         |          ✅         |       |see [Reference Commands](vllm\u002FREADME.md\u002F#33-reference-commands-for-running-gemma-4-models-and-diffusiongemma)                            |\n| THUDM\u002FGLM-4v-9B                            |  ✅  |         ✅         |          ✅         |       |  with --hf-overrides and chat_template  |\n| zai-org\u002FGLM-4.1V-9B-Base                   |  ✅  |         ✅         |          ✅          |       |                           |\n| zai-org\u002FGLM-4.1V-9B-Thinking               |  ✅  |         ✅         |          ✅          |       |                           |\n| zai-org\u002FGlyph                              |  ✅  |         ✅         |          ✅          |       |                           |\n| opendatalab\u002FMinerU2.5-2509-1.2B            |  ✅  |         ✅         |          ✅          |       |                           |\n| baidu\u002FERNIE-4.5-VL-28B-A3B-Thinking        |  ✅  |         ✅         |          ✅          |       |                           |\n| zai-org\u002FGLM-4.6V-Flash                     |  ✅  |         ✅         |          ✅          |       |   pip install transformers==5.0.0rc0 first            |\n| PaddlePaddle\u002FPaddleOCR-VL                  |  ✅  |         ✅         |          ✅          |       |  follow the guide in [here](.\u002Fvllm\u002FREADME.md#32-how-to-use-paddleocr)     |\n| deepseek-ai\u002FDeepSeek-OCR                   |  ✅  |         ✅         |          ✅          |       |                           |\n| deepseek-ai\u002FDeepSeek-OCR-2                 |  ✅  |         ✅         |          ✅          |       |  There may be accuracy issues when using `--quantization fp8`             |\n| moonshotai\u002FKimi-VL-A3B-Thinking-2506       |  ✅  |         ✅         |          ✅          |       |                           |\n| Qwen\u002FQwen2.5-Omni-7B                       |  ✅  |         ✅         |          ✅          |       |                           |\n| Qwen\u002FQwen3-Omni-30B-A3B-Instruct           |  ✅  |         ✅         |          ✅          |       |                           |\n| openai\u002Fwhisper-medium                      |  ✅  |         ✅         |          ✅          |       |                           |\n| openai\u002Fwhisper-large-v3                    |  ✅  |         ✅         |          ✅          |       |                           |\n| Qwen\u002FQwen3-Embedding-8B                    |  ✅  |         ✅         |          ✅          |       |                           |\n| Qwen3-VL-Embedding-2B\u002F8B                   |  ✅  |         ✅         |          ✅          |       |  follow the guide in [here](https:\u002F\u002Fgithub.com\u002Fvllm-project\u002Fvllm\u002Fblob\u002F2f4226fe5280b60c47b4f6f01d9b18ac9cda2038\u002Fexamples\u002Fpooling\u002Fembed\u002Fvision_embedding_online.py)                    |\n| BAAI\u002Fbge-m3                                |  ✅  |         ✅         |          ✅          |       |                           |\n| BAAI\u002Fbge-large-en-v1.5                     |  ✅  |         ✅         |          ✅          |       |                           |\n| Qwen\u002FQwen3-Reranker-8B                     |  ✅  |         ✅         |          ✅          |       |                           |\n| Qwen3-VL-Reranker-2B\u002F8B                    |  ✅  |         ✅         |          ✅          |       |  follow the guide in [here](https:\u002F\u002Fgithub.com\u002Fvllm-project\u002Fvllm\u002Fblob\u002F2f4226fe5280b60c47b4f6f01d9b18ac9cda2038\u002Fexamples\u002Fpooling\u002Fscore\u002Fvision_rerank_api_online.py)                    |\n| BAAI\u002Fbge-reranker-large                    |  ✅  |         ✅         |          ✅          |       |                           |\n| BAAI\u002Fbge-reranker-v2-m3                    |  ✅  |         ✅         |          ✅          |       |                           |\n\n\n\n--- \n\n\n## LLM Scaler Omni (experimental)\n\n`llm-scaler-omni` supports running image\u002Fvoice\u002Fvideo generation etc., featuring `Omni Studio` mode (using ComfyUI) and `Omni Serving` mode (via SGLang Diffusion or Xinference).  \n\n\nPlease follow the instructions in the [Getting Started](omni\u002FREADME.md#getting-started-with-the-omni-docker-image) to use `llm-scaler-omni`.\n\n\n### Omni Demos\n\n| Qwen-Image | Multi B60 Wan2.2-T2V-14B |\n|------------|--------------------------|\n| ![Qwen Image Demo](.\u002Fomni\u002Fassets\u002Fdemo_qwen_image.gif) | ![Wan2.2 T2V Demo](.\u002Fomni\u002Fassets\u002Fdemo_wan2.2_14b_i2v_multi_xpu.gif) |\n\n\n### Omni Studio (ComfyUI WebUI interaction)\n\n`Omni Stuido` supports Image Generation\u002FEdit, Video Generation, Audio Generation, 3D Generation etc.  \n\n\n| Model Category | Model | Type | \n|----------------------|------------|---------------|\n| **Image Generation** | Qwen-Image, Qwen-Image-Edit | Text-to-Image, Image Editing | \n| **Image Generation** | Stable Diffusion 3.5 | Text-to-Image, ControlNet | \n| **Image Generation** | Z-Image-Turbo | Text-to-Image | \n| **Image Generation** | Flux.1, Flux.1 Kontext dev | Text-to-Image, Multi-Image Reference, ControlNet | \n| **Image Generation** | FireRed-Image-Edit-1.1 | Image Editing | \n| **Video Generation** | Wan2.2 TI2V 5B, Wan2.2 T2V 14B, Wan2.2 I2V 14B | Text-to-Video, Image-to-Video | \n| **Video Generation** | Wan2.2 Animate 14B | Video Animation | \n| **Video Generation** | HunyuanVideo 1.5 8.3B | Text-to-Video, Image-to-Video | \n| **Video Generation** | LTX-2 | Text-to-Video, Image-to-Video | \n| **3D Generation** | Hunyuan3D 2.1 | Text\u002FImage-to-3D | \n| **Audio Generation** | VoxCPM1.5, IndexTTS 2 | Text-to-Speech, Voice Cloning | \n| **Video Upscaling** | SeedVR2 | Video Restoration and Upscaling | \n\n\nPlease check [ComfyUI Support](omni\u002Fdocs\u002FCOMFYUI.md) for more details.\n\n### Omni Serving (OpenAI-API compatible serving)\n\n`Omni Serving` supports Image Generation, Audio Generation etc.\n\n- Image Generation (`\u002Fv1\u002Fimages\u002Fgenerations`): Stable Diffusion 3.5, Flux.1-dev\n- Text to Speech (`\u002Fv1\u002Faudio\u002Fspeech`): Kokoro 82M\n- Speech to Text (`\u002Fv1\u002Faudio\u002Ftranscriptions`): whisper-large-v3\n\nPlease check the [b8 Xinference documentation](https:\u002F\u002Fgithub.com\u002Fintel\u002Fllm-scaler\u002Fblob\u002Fomni-0.1.0-b8\u002Fomni\u002FREADME.md#xinference) for more details.\n\n---\n## Releases\n- Please check out the Docker image releases for [llm-scaler-vllm](Releases.md\u002F#llm-scaler-vllm) and [llm-scaler-omni](Releases.md\u002F#llm-scaler-omni)\n\n---\n## Get Support\n- Please report a bug or raise a feature request by opening a [Github Issue](https:\u002F\u002Fgithub.com\u002Fintel\u002Fllm-scaler\u002Fissues)\n","LLM Scaler 是面向 Intel® Arc™ Pro B60\u002FB70 GPU 的生成式 AI 推理优化框架，专为高效运行大语言模型（LLM）与多模态生成模型（文本、图像、视频）而设计。它基于 vLLM、ComfyUI、SGLang Diffusion 和 Xinference 等主流框架进行深度适配，支持 FP8\u002FINT4 量化、多 Token 预测（MTP）、LoRA 服务、XPU 图加速及模型流式加载等关键技术，显著提升吞吐量与首字延迟（TTFT）。适用于在 Intel 独立显卡上部署本地化 GenAI 应用的开发者与企业，尤其适合需要兼顾性能、精度与硬件兼容性的边缘或工作站级生成任务。",2,"2026-08-21 02:30:05","trending"]