[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-92699":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":9,"htmlUrl":10,"language":11,"languages":10,"totalLinesOfCode":10,"stars":12,"forks":13,"watchers":14,"openIssues":13,"contributorsCount":13,"subscribersCount":13,"size":13,"stars1d":13,"stars7d":13,"stars30d":15,"stars90d":13,"forks30d":13,"starsTrendScore":13,"compositeScore":16,"rankGlobal":10,"rankLanguage":10,"license":17,"archived":18,"fork":18,"defaultBranch":19,"hasWiki":20,"hasPages":18,"topics":21,"createdAt":10,"pushedAt":10,"updatedAt":22,"readmeContent":23,"aiSummary":24,"trendingCount":13,"starSnapshotCount":13,"syncStatus":25,"lastSyncTime":26,"discoverSource":27},92699,"DyRef","Weistrass\u002FDyRef","Weistrass","Official repository for Scaling Multi-Reference Image Generation with Dynamic Reward Optimization (ECCV2026)","",null,"Python",63,0,1,7,40.7,"Apache License 2.0",false,"main",true,[],"2026-07-22 04:02:06","\u003Cdiv align=\"center\">\n\n# Scaling Multi-Reference Image Generation with Dynamic Reward Optimization\n\nWenwang Huang\u003Csup>1,\\*\u003C\u002Fsup>, Yusen Fu\u003Csup>1,\\*\u003C\u002Fsup>, Junjie Wang\u003Csup>1\u003C\u002Fsup>, Mengfei Huang\u003Csup>1\u003C\u002Fsup>, Yulin Li\u003Csup>1\u003C\u002Fsup>, Gan Liu\u003Csup>2\u003C\u002Fsup>, Jing Cai\u003Csup>2\u003C\u002Fsup>, Yancheng He\u003Csup>2\u003C\u002Fsup>, Zhuotao Tian\u003Csup>1,3,†\u003C\u002Fsup>\n\n\u003Csup>1\u003C\u002Fsup> Harbin Institute of Technology, Shenzhen,  \n\u003Csup>2\u003C\u002Fsup> Independent Researcher,  \n\u003Csup>3\u003C\u002Fsup> Shenzhen Loop Area Institute,  \n\u003Csup>*\u003C\u002Fsup> Equal contribution · \u003Csup>†\u003C\u002Fsup> Corresponding author\n\n[![ECCV](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FECCV-2026-8bc34a?style=flat&labelColor=555555)](https:\u002F\u002Feccv.ecva.net\u002F)\n[![Paper](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FPaper-arXiv-b23a2f?style=flat&labelColor=555555)](https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.26947)\n[![License](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FLicense-Apache%202.0-d4b31a?style=flat&labelColor=555555)](LICENSE)\n[![Project](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FProject-Page-63b32e?style=flat&labelColor=555555)](#results-gallery)\n[![Benchmark](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FBenchmark-OmniRef--Bench-c0362c?style=flat&labelColor=555555)](https:\u002F\u002Fhuggingface.co\u002Fdatasets\u002FEason0438\u002FOmniRef-Bench)\n[![Dataset](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FDataset-Training-3f3f3f?style=flat&labelColor=555555)](https:\u002F\u002Fhuggingface.co\u002Fdatasets\u002FEason0438\u002FOmniRef-training)\n[![Model](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FModel-Weights-4c1?style=flat&labelColor=555555)](https:\u002F\u002Fhuggingface.co\u002FWeistrass\u002FQwen-Image-Edit-2511-DyRef)\n\n\u003C\u002Fdiv>\n\n## 🔖 Table of Contents\n\n1. [News](#news)\n2. [Todo List](#todo-list)\n3. [Highlights](#highlights)\n4. [Results Gallery](#results-gallery)\n5. [Motivation](#motivation)\n6. [Method](#method)\n7. [Installation](#installation)\n8. [Quickstart](#quickstart)\n9. [Evaluation](#evaluation)\n10. [Acknowledgement](#acknowledgement)\n11. [Citation](#citation)\n\u003C!-- 12. [Star History](#star-history) -->\n\n## 🔥News\n\n- [2026.06.18] Our paper has been accepted by ECCV 2026 🎉🎉🎉. Congratulations to all collaborators! 🎊🎊🎊\n- [2026.07.02] We have released the training and inference code of DyRef, along with the corresponding training data, benchmark, and model weights. Feel free to use it! ⭐⭐⭐\n\n## 📋Todo List\n\n- [x] Release our paper on arXiv.\n- [x] Open-source the training code of DyRef for Qwen-Image-Edit-2511, and Flux2-klein-base.\n- [x] Release the training data used in our work.\n- [x] Open-source OmniRef-Bench, our proposed benchmark.\n- [x] Release the model weights trained by our method.\n- [ ] Improve code efficiency and structure.\n- [ ] Release the project page\n\n## ✨Highlights\n\n\u003Cp align=\"center\">\n  \u003Cimg src=\"assets\u002Ffigure1.png\" alt=\"DyRef Performance\" \u002F>\n\u003C\u002Fp>\n\n![DyRef Quantitative](assets\u002Fradar_plot.png)\n\n1. We introduce [**OmniRef-Bench**](https:\u002F\u002Fhuggingface.co\u002Fdatasets\u002FEason0438\u002FOmniRef-Bench) to evaluate complex multi-reference image generation (MRIG) across diverse combinations of reference-image types and quantities. Our evaluation reveals that current open-source models still struggle substantially on this challenging benchmark.\n\n2. We develop an automated data synthesis pipeline for complex MRIG and propose **DyRef** to address the sharp performance degradation of existing open-source models as reference complexity increases. Trained with DyRef, Qwen-Image-Edit-2511 achieves performance comparable to the closed-source model **Nano Banana Pro** and surpasses **Seedream 4.5**.\n\n3. DyRef consistently enhances open-source models, including Qwen-Image-Edit-2511 and FLUX.2-klein-base, on MRIG benchmarks (**OmniRef-Bench**, **MultiBanana**, and **OmniContext**) while also improving single-image editing performance on **ImgEdit** and **DreamBench++**, without compromising the foundation models' other capabilities.\n\n## 🖼️Results Gallery\nIn this part, we present qualitative comparisons between our method and existing approaches across various combinations of reference types. The prompts shown in the figures are abbreviated versions for visualization purposes and are not the actual prompts used during inference. In addition, we use color to indicate the attributes specified in the user prompt and their corresponding reference images.\n\n![case1](assets\u002Fsuccess_cases_p3_1_cropped-1.png)\n\n![case2](assets\u002Fsuccess_cases_p3_2_cropped-1.png)\n\n## 💡Motivation\n![Motivation](assets\u002Fmotivation.png)\n\nIn this work, we observe that the performance of current mainstream open-source image generation models degrades substantially as the number of reference images increases and the diversity of reference types expands. This limitation significantly restricts their potential for professional applications, such as creative image generation and video keyframe generation.\n\n\n## 🧩Method\n![Method](assets\u002Fmethod.png)\n\n**Illustration of DyRef**. To address this issue, we propose DyRef, a two-stage training framework. \n\nIn Stage I, we employ SFT to equip the model with the basic capability to handle complex MRIG tasks. In Stage II, **(a) DRS** enlarges the reward differences across samples for better training, and **(b) DAR** enhances the model’s focus on samples with numerous mixed-type reference images.\n\n## ⚙️Installation\n\nDyRef uses separate Conda environments for SFT, RL, and benchmark evaluation because each component has its own dependency requirements.\n\n### SFT Environment\n\n```bash\nconda create -n dyref_sft python=3.11 -y\nconda activate dyref_sft\ncd DyRef\u002Fsft\npip install -e .\n```\n\n### RL Environment\n\n```bash\nconda create -n dyref_rl python=3.11 -y\nconda activate dyref_rl\ncd DyRef\u002Frl\npip install -e .[deepspeed]\n```\n\n### Model Download\nPlease download the following models from Hugging Face: \n- Qwen\u002FQwen-Image-Edit-2511\n- google\u002Fsiglip2-base-patch16-384\n- facebook\u002Fdinov2-base\n- openai\u002Fclip-vit-large-patch14\n- black-forest-labs\u002FFLUX.2-klein-base-9B (if you need)\n\nFor example:\n```bash\nhf download Qwen\u002FQwen-Image-Edit-2511 --local-dir \"your path to save this model\"\n```\n\n## 🚀Quickstart\n\n### Model Inference\nBefore running the following command, please download the [Qwen-Image-Edit-2511](https:\u002F\u002Fhuggingface.co\u002FQwen\u002FQwen-Image-Edit-2511) model and the [pretrained LoRA weights](https:\u002F\u002Fhuggingface.co\u002FWeistrass\u002FQwen-Image-Edit-2511-DyRef) we provide.\n\n```bash\nconda activate dyref_sft\npython model_inference\u002FQwen-Image-Edit-2511.py\n```\n\n### Train\nIf you would like to train a model from scratch using your own data or reproduce our training pipeline, please run the following command.\n\n#### 1. Stage 1: SFT Training\n\n```bash\nconda activate dyref_sft\ncd DyRef\u002Fsft\nbash all_scripts\u002FQwen-Image-Edit-2511_lora.sh\n```\n\n#### 2. Convert SFT Weights for RL\nSince the original repositories used for SFT and RL training adopt different formats for storing LoRA weights, the following conversion script is required to ensure that the LoRA weights generated during the SFT stage can be correctly loaded during RL training.\n```bash\ncd DyRef\u002Fsft\npython all_scripts\u002Fdiffusers_peft_transfer.py --mode d2p \\\n    --input \u002Fpath\u002Fto\u002Fsft_checkpoint.safetensors \\\n    --output \u002Fpath\u002Fto\u002Foutput_peft_format_dir \\\n    --prefix transformer \\\n    --verify\n```\n\n#### 3. Stage 2: RL Training\n\n```bash\nconda activate dyref_rl\ncd DyRef\u002Frl\nbash scripts\u002Fqwen2511-gdpo-rank64-add2k5-csd-siglipv2_flat-sigmoid0.65-focal_loss.sh\n```\n\n## 📊Evaluation\n\nOmniRef-Bench is designed to measure whether generated images preserve and combine multiple references in a balanced way.\n\nThe benchmark covers:\n- Subject fidelity\n- Style consistency\n- Background consistency\n- Lighting consistency\n- Pose consistency\n- Overall multi-reference alignment\n\nThe evaluation-related code is organized under:\n- `benchmark\u002FGrounded-SAM-2_patch\u002F`\n- `benchmark\u002FCSD_patch\u002F`\n- `benchmark\u002FAlphaPose_patch\u002F`\n- `benchmark\u002FMLLM_eval\u002F`\n\n### 1. Convert RL Weights for Evaluation\n\n```bash\ncd DyRef\u002Fsft\npython all_scripts\u002Fdiffusers_peft_transfer.py --mode p2d \\\n    --input \u002Fpath\u002Fto\u002Frl_checkpoint_dir \\\n    --output \u002Fpath\u002Fto\u002Foutput.safetensors \\\n    --prefix '' \\\n    --verify\n```\n\n### 2. Generate Images for Evaluation\n\n```bash\nconda activate dyref_sft\ncd DyRef\u002Fsft\nbash all_scripts\u002Feval\u002Feval_ourbench_lora_2511.sh\n```\n\n### 3. Run OmniRef-Bench\n\n```bash\ncd benchmark\nbash eval_suite.sh \\\n    \u002Fpath\u002Fto\u002Fgenerated_images \\\n    \u002Fpath\u002Fto\u002Foutput_dir \\\n    \u002Fpath\u002Fto\u002Ftest_set \\\n    \u002Fpath\u002Fto\u002Ftest_set.json\n```\n\nFor more details, see [benchmark\u002FREADME.md](benchmark\u002FREADME.md).\n\n## 🙏Acknowledgement\n\nDyRef is built on several excellent open-source projects:\n- [DiffSynth-Studio](https:\u002F\u002Fgithub.com\u002Fmodelscope\u002Fdiffsynth-studio) for SFT training infrastructure\n- [Flow-Factory](https:\u002F\u002Fgithub.com\u002FX-GenGroup\u002FFlow-Factory) for RL training infrastructure\n- [Grounded-SAM-2](https:\u002F\u002Fgithub.com\u002FIDEA-Research\u002FGrounded-SAM-2) for subject and background evaluation\n- [CSD](https:\u002F\u002Fgithub.com\u002Flearn2phoenix\u002FCSD) for style consistency evaluation\n- [AlphaPose](https:\u002F\u002Fgithub.com\u002FMVIG-SJTU\u002FAlphaPose) for pose evaluation\n\n## 📝Citation\n\nIf you find DyRef useful in your research, please consider citing it:\n\n```bibtex\n@article{huang2026scaling,\n  title={Scaling Multi-Reference Image Generation with Dynamic Reward Optimization},\n  author={Huang, Wenwang and Fu, Yusen and Wang, Junjie and Huang, Mengfei and Li, Yulin and Liu, Gan and Cai, Jing and He, Yancheng and Tian, Zhuotao},\n  journal={arXiv preprint arXiv:2606.26947},\n  year={2026}\n}\n```\n\n\u003C!-- \u003Ca id=\"star-history\">\u003C\u002Fa>\n\n## ⭐️Star History\n\n[![Star History Chart](https:\u002F\u002Fapi.star-history.com\u002Fsvg?repos=Weistrass\u002FDyRef&type=Date)](https:\u002F\u002Fstar-history.com\u002F#Weistrass\u002FDyRef&Date) -->\n","DyRef 是一个面向多参考图像生成（MRIG）任务的动态奖励优化框架，旨在提升模型在融合多个参考图像（如风格图、结构图、语义图等）时的生成质量与一致性。其核心通过可学习的动态奖励机制调节不同参考信号的贡献权重，并支持主流多模态基础模型（如Qwen-Image-Edit、Flux2）的微调与推理。项目配套发布专用评测基准OmniRef-Bench、训练数据集及预训练权重，强调对复杂参考组合（跨类型、多数量）的泛化能力。适用于图像编辑、可控图像生成、跨模态内容合成等需精细参考引导的视觉生成场景。",2,"2026-07-10 02:30:12","CREATED_QUERY"]