[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-95934":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":8,"htmlUrl":8,"language":9,"languages":8,"totalLinesOfCode":8,"stars":10,"forks":11,"watchers":12,"openIssues":12,"contributorsCount":13,"subscribersCount":13,"size":13,"stars1d":13,"stars7d":13,"stars30d":14,"stars90d":13,"forks30d":13,"starsTrendScore":13,"compositeScore":15,"rankGlobal":8,"rankLanguage":8,"license":8,"archived":16,"fork":16,"defaultBranch":17,"hasWiki":16,"hasPages":16,"topics":18,"createdAt":8,"pushedAt":8,"updatedAt":19,"readmeContent":20,"aiSummary":21,"trendingCount":13,"starSnapshotCount":13,"syncStatus":22,"lastSyncTime":23,"discoverSource":24},95934,"LLaDA-Image","inclusionAI\u002FLLaDA-Image","inclusionAI",null,"Python",130,6,1,0,26,2.54,false,"main",[],"2026-09-21 02:04:29","\u003Ch1 align=\"center\">LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes\u003C\u002Fh1>\n\n\u003Cp align=\"center\">\n  Welcome to the official repository for LLaDA-Image, a unified model for high-quality image generation and editing.\n\u003C\u002Fp>\n\n\u003Cp align=\"center\">\n  \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FinclusionAI\u002FLLaDA-Image\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FGitHub-LLaDA--Image-181717?logo=github\" alt=\"GitHub\">\u003C\u002Fa>\n    \u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fpdf\u002F2609.03796\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FarXiv-Report-B31B1B?logo=arxiv\" alt=\"arXiv\">\u003C\u002Fa>\u003Cbr>\n  \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002FinclusionAI\u002FLLaDA-Image\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FHugging%20Face-Base-FFD21E?logo=huggingface\" alt=\"LLaDA-Image Base on Hugging Face\">\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002FinclusionAI\u002FLLaDA-Image-FP8\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FHugging%20Face-Base--FP8-FFD21E?logo=huggingface\" alt=\"LLaDA-Image Base FP8 Version on Hugging Face\">\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002FinclusionAI\u002FLLaDA-Image-Turbo\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FHugging%20Face-Turbo-FFD21E?logo=huggingface\" alt=\"LLaDA-Image Turbo on Hugging Face\">\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002FinclusionAI\u002FLLaDA-Image-Turbo-FP8\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FHugging%20Face-Turbo--FP8-FFD21E?logo=huggingface\" alt=\"LLaDA-Image Turbo FP8 Version on Hugging Face\">\u003C\u002Fa>\u003Cbr>\n    \u003Ca href=\"https:\u002F\u002Fmodelscope.cn\u002Fmodels\u002FinclusionAI\u002FLLaDA-Image\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002F🤖%20Model%20Scope-Base-624aff\" alt=\"LLaDA-Image Base on Modelscope\">\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fmodelscope.cn\u002Fmodels\u002FinclusionAI\u002FLLaDA-Image-FP8\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002F🤖%20Model%20Scope-Base--FP8-624aff\" alt=\"LLaDA-Image Base FP8 Version on Modelscope\">\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fmodelscope.cn\u002Fmodels\u002FinclusionAI\u002FLLaDA-Image-Turbo\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002F🤖%20Model%20Scope-Turbo-624aff\" alt=\"LLaDA-Image Turbo on Modelscope\">\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fmodelscope.cn\u002Fmodels\u002FinclusionAI\u002FLLaDA-Image-Turbo-FP8\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002F🤖%20Model%20Scope-Turbo--FP8-624aff\" alt=\"LLaDA-Image Turbo FP8 Version on Modelscope\">\u003C\u002Fa>\n\u003C\u002Fp>\n\n\u003Cp align=\"center\">\n  \u003Cimg src=\".\u002Fassets\u002Fdemo.jpg\" alt=\"LLaDA-Image realistic image generation showcase\" width=\"100%\">\n  \u003Cbr>\n  \u003Cem>Photorealistic image generation with natural lighting, lifelike details, and coherent scenes.\u003C\u002Fem>\n\u003C\u002Fp>\n\n\u003Cp align=\"center\">\n  \u003Cimg src=\".\u002Fassets\u002Fdemo_p2.jpg\" alt=\"LLaDA-Image text rendering and poster generation showcase\" width=\"100%\">\n  \u003Cbr>\n  \u003Cem>High-quality text rendering and creative poster generation across diverse visual styles.\u003C\u002Fem>\n\u003C\u002Fp>\n\n\u003Cp align=\"center\">\n  \u003Cimg src=\".\u002Fassets\u002Fedit_demo.jpg\" alt=\"LLaDA-Image editing showcase\" width=\"100%\">\n  \u003Cbr>\n  \u003Cem>Instruction-guided image editing with faithful content preservation and precise visual changes.\u003C\u002Fem>\n\u003C\u002Fp>\n\n## Introduction\n\nLLaDA-Image is a competitive 6B-parameter open-source unified image generation and editing model family. It includes **LLaDA-Image**, a 50-step Base model for high-quality text-to-image generation and instruction-guided editing, and **LLaDA-Image-Turbo**, a 4-step distilled model for fast generation and editing. Both variants support practical text-to-image generation, VQ-conditioned generation, reference-image editing, and Chinese--English text rendering.\n\nThis repository provides the checkpoints and Diffusers-based inference code for the LLaDA-Image model family.\n\n## News\n\n- **2026-09-04:** We released the LLaDA-Image Base and Turbo checkpoints together with the inference code.\n\n## Highlights\n\n- **Unified generation and editing.** A single checkpoint supports text-to-image generation and reference-preserving, instruction-guided editing without a separate editing backbone.\n- **Unified diffusion model.** Both backbone and DiT are diffusion models, trained in a unified framework.\n- **Realistic image generation.** LLaDA-Image produces high-quality images with rich visual details, natural lighting, and coherent compositions.\n- **Image-only pre-training for visual-prior learning.** The report establishes the visual prior through image-only pre-training and mid-training before introducing paired language supervision and joint generation--editing training.\n- **Efficient inference with distilled model.** LLaDA-Image-Turbo uses Twin-DMD distillation to deliver fast image generation and editing in only 2--4 sampling steps.\n- **SOTA on Qwen-Image-Bench.** LLaDA-Image achieves state-of-the-art overall scores of 53.53 in English and 53.38 in Chinese.\n\n\u003Cp align=\"center\">\n  \u003Cimg class=\"not-prose\" src=\".\u002Fassets\u002Fqwen-imagebench.png\" alt=\"Qwen-image bench evaluation\" width=\"100%\">\n\u003C\u002Fp>\n\n## Model Zoo\n\n| Model                 | Description                                                                           | Sampling steps | Hugging Face (Checkpoints)                                                                    | ModelScope (Checkpoints)                                                                          |\n| --------------------- | ------------------------------------------------------------------------------------- | -------------: | --------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------- |\n| **LLaDA-Image**       | Base model for high-fidelity text-to-image generation and instruction-guided editing. |             50 | **BF16:** [inclusionAI\u002FLLaDA-Image](https:\u002F\u002Fhuggingface.co\u002FinclusionAI\u002FLLaDA-Image)\u003Cbr>**FP8:** [inclusionAI\u002FLLaDA-Image-FP8](https:\u002F\u002Fhuggingface.co\u002FinclusionAI\u002FLLaDA-Image-FP8) | **BF16:** [inclusionAI\u002FLLaDA-Image](https:\u002F\u002Fmodelscope.cn\u002Fmodels\u002FinclusionAI\u002FLLaDA-Image)\u003Cbr>**FP8:** [inclusionAI\u002FLLaDA-Image-FP8](https:\u002F\u002Fmodelscope.cn\u002Fmodels\u002FinclusionAI\u002FLLaDA-Image-FP8) |\n| **LLaDA-Image-Turbo** | Distilled model for fast generation and editing.                                      |              4 | **BF16:** [inclusionAI\u002FLLaDA-Image-Turbo](https:\u002F\u002Fhuggingface.co\u002FinclusionAI\u002FLLaDA-Image-Turbo)\u003Cbr>**FP8:** [inclusionAI\u002FLLaDA-Image-Turbo-FP8](https:\u002F\u002Fhuggingface.co\u002FinclusionAI\u002FLLaDA-Image-Turbo-FP8) | **BF16:** [inclusionAI\u002FLLaDA-Image-Turbo](https:\u002F\u002Fmodelscope.cn\u002Fmodels\u002FinclusionAI\u002FLLaDA-Image-Turbo)\u003Cbr>**FP8:** [inclusionAI\u002FLLaDA-Image-Turbo-FP8](https:\u002F\u002Fmodelscope.cn\u002Fmodels\u002FinclusionAI\u002FLLaDA-Image-Turbo-FP8) |\n\n## Opensource Plan\n\n- [x] Inference code and model weights\n- [ ] Training code (coming soon)\n\n## Quick Start\n\n### 1. Create an environment\n\nThe implementation has been used with Python 3.11, PyTorch 2.8, Transformers 4.57.6, and Diffusers 0.39.0.\n\n```bash\ngit clone https:\u002F\u002Fgithub.com\u002FinclusionAI\u002FLLaDA-Image.git\ncd LLaDA-Image\n\nconda create -n llada-image python=3.11 -y\nconda activate llada-image\n\npip install -r requirements.txt\n```\n\n### 2. Run inference\n\nThe pipeline accepts a prompt and, for editing, an optional reference image.\n\n#### LLaDA-Image (Base)\n\nUse the Base checkpoint for high-fidelity generation and editing. Its recommended sampling configuration is **50 steps**.\n\n```python\nimport torch\n\nfrom src import LLaDAImagePipeline\n\n# Load the pipeline. The model is downloaded from Hugging Face on first use.\npipe = LLaDAImagePipeline.from_pretrained(\n    \"inclusionAI\u002FLLaDA-Image\",\n    torch_dtype=torch.bfloat16,\n    device=\"cuda\",\n)\n\n# Generate an image.\nprompt = (\n    \"A cinematic photograph of a red fox standing in fresh snow, \"\n    \"soft winter light, detailed fur, shallow depth of field\"\n)\nnegative_prompt = \"\"\n\nimage = pipe(\n    prompt=prompt,\n    negative_prompt=negative_prompt,\n    generation_mode=\"text\",\n    height=1024,\n    width=1024,\n    num_inference_steps=50,\n    guidance_scale=5.0,\n    generator=torch.Generator(\"cuda\").manual_seed(42),\n).images[0]\n\nimage.save(\"llada-image-base.png\")\n```\n\n#### LLaDA-Image-Turbo\n\nUse the Turbo checkpoint for fast generation and editing. Its recommended sampling configuration is **4 steps**.\n\n> [!NOTE]\n> For LLaDA-Image-Turbo inference, you can try setting `stochastic_sampling` to `false` in `scheduler\u002Fscheduler_config.json`, which may produce sharper details in some cases.\n\n```python\nimport torch\n\nfrom src import LLaDAImagePipeline\n\n# Load the distilled Turbo checkpoint.\npipe = LLaDAImagePipeline.from_pretrained(\n    \"inclusionAI\u002FLLaDA-Image-Turbo\",\n    torch_dtype=torch.bfloat16,\n    device=\"cuda\",\n)\n\nprompt = \"A quiet observatory above a sea of clouds at sunrise, golden light, wide-angle photograph\"\n\nimage = pipe(\n    prompt=prompt,\n    generation_mode=\"text\",\n    height=1024,\n    width=1024,\n    num_inference_steps=4,\n    guidance_scale=1.0,\n    generator=torch.Generator(\"cuda\").manual_seed(42),\n).images[0]\n\nimage.save(\"llada-image-turbo.png\")\n```\n\n#### Generation modes\n\nBoth checkpoints support the following modes. Text and VQ-conditioned generation require height and width divisible by 16; image editing requires dimensions divisible by 32.\n\n**VQ-conditioned generation** uses the LLaDA2 model to produce image VQ tokens from the prompt, which SigVQ embeds before diffusion. Do not provide an input image in VQ mode.\n\n```python\nimage = pipe(\n    prompt=\"A quiet observatory above a sea of clouds at sunrise\",\n    generation_mode=\"vq\",\n    height=1024,\n    width=1024,\n    num_inference_steps=50,  # Use 4 for LLaDA-Image-Turbo.\n    guidance_scale=5.0,  # Use 1.0 for few-step inference.\n    generator=torch.Generator(\"cuda\").manual_seed(42),\n).images[0]\n```\n\n**Image editing** requires a reference image:\n\n```python\nfrom diffusers.utils import load_image\n\nreference_image = load_image(\"\u002Fpath\u002Fto\u002Finput.png\")\nimage = pipe(\n    prompt=\"Turn it into a watercolor painting\",\n    image=reference_image,\n    generation_mode=\"editing\",\n    height=1024,\n    width=1024,\n    num_inference_steps=50,  # Use 4 for LLaDA-Image-Turbo.\n    guidance_scale=5.0,  # Use 1.0 for few-step inference.\n    generator=torch.Generator(\"cuda\").manual_seed(43),\n).images[0]\n```\n\n## Acknowledgements\n\nWe thank the [VeOmni](https:\u002F\u002Fgithub.com\u002FByteDance-Seed\u002FVeOmni) project and its contributors for their valuable open-source work.\n\n## Citation\n\nIf you find LLaDA-Image useful for your research or applications, please consider citing our work.\n\n```bibtex\n@article{LLaDAImage,\ntitle = {LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes},\nauthor = {Chuyan Chen and Haoxing Chen and Kun Chen and Zhenglin Cheng and Long Cui and Ruishan Fang and Zhangxuan Gu and Zhicheng Huang and Zhenzhong Lan and Yuanting Lei and Haoquan Li and Jianguo Li and Rongchuan Li and Sidu Li and Tao Lin and Deyuan Liu and Jiacheng Liu and Lin Liu and Yuxuan Lou and Zhisheng Lu and Yuxin Ma and Shuheng Shen and Peng Sun and Chaoyang Wang and Hongjun Wang and Xiaomei Wang and Yongxin Wang and Chengzhang Wu and Hongru Wu and Jun Xie},\njournal = {arXiv preprint arXiv:2609.03796},\nyear = {2026}\n}\n```\n","LLaDA-Image 是一个开源的统一图像生成与编辑模型家族，支持高质量文生图、指令驱动图像编辑及多风格文本渲染。其核心基于6B参数量架构，提供Base\u002FTurbo双版本及FP8量化变体，强调完全公开的训练配方与可复现性。模型在写实光影、细节保真、场景连贯性及文本可读性方面表现突出，适用于AIGC内容创作、设计辅助、营销素材生成等需兼顾生成质量与可控编辑能力的视觉生产场景。",2,"2026-09-06 02:30:06","CREATED_QUERY"]