[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-93059":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":9,"htmlUrl":10,"language":10,"languages":10,"totalLinesOfCode":10,"stars":11,"forks":12,"watchers":13,"openIssues":14,"contributorsCount":15,"subscribersCount":15,"size":15,"stars1d":15,"stars7d":15,"stars30d":15,"stars90d":15,"forks30d":15,"starsTrendScore":15,"compositeScore":16,"rankGlobal":10,"rankLanguage":10,"license":10,"archived":17,"fork":17,"defaultBranch":18,"hasWiki":19,"hasPages":17,"topics":20,"createdAt":10,"pushedAt":10,"updatedAt":21,"readmeContent":22,"aiSummary":23,"trendingCount":15,"starSnapshotCount":15,"syncStatus":12,"lastSyncTime":24,"discoverSource":25},93059,"PixWorld","SensenGao\u002FPixWorld","SensenGao","PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space","",null,138,2,17,3,0,38.43,false,"main",true,[],"2026-07-22 04:02:08","\u003Cdiv align=\"center\">\n\u003Ch1>\nPixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space\u003C\u002Fh1>\n\n[Sensen Gao\\*](https:\u002F\u002Fsensengao.github.io\u002F)\u003Csup>1\u003C\u002Fsup>, [Zhaoqing Wang\\*](https:\u002F\u002Fderrickwang005.github.io\u002F)\u003Csup>2\u003C\u002Fsup>, [Qihang Cao](https:\u002F\u002Fscholar.google.com\u002Fcitations?user=oegbT6AAAAAJ&hl=zh-CN)\u003Csup>1\u003C\u002Fsup>, [Dongdong Yu](https:\u002F\u002Fscholar.google.com\u002Fcitations?user=B2RmjSYAAAAJ&hl=zh-CN)\u003Csup>2\u003C\u002Fsup>, [Changhu Wang](https:\u002F\u002Fscholar.google.com\u002Fcitations?user=DsVZkjAAAAAJ&hl=en)\u003Csup>2\u003C\u002Fsup>, [Jia-Wang Bian📧](https:\u002F\u002Fjwbian.net\u002F)\u003Csup>1\u003C\u002Fsup>\n\n\u003Csup>1\u003C\u002Fsup> Nanyang Technological University &nbsp;&nbsp; \u003Csup>2\u003C\u002Fsup> AISphere &nbsp;&nbsp;|&nbsp;&nbsp; \\* Co-first authors &nbsp;&nbsp; 📧 Corresponding author\n\n\u003Ca href=\"https:\u002F\u002Fsensengao.github.io\u002FPixWorld\u002F\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FProject_Page-yellowgreen\" alt=\"Project Page\">\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.05373\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FarXiv-2607.05373-b31b1b\" alt=\"arXiv\">\u003C\u002Fa>\n\u003Ca href=\"#-todo\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FWeights-PixWorld--480P--4steps_coming_soon-blue\" alt=\"Weights\">\u003C\u002Fa>\n\n\u003Cp align=\"center\">\n  \u003Ca href=\"https:\u002F\u002Fsensengao.github.io\u002FPixWorld\u002F\">\n    \u003Cimg src=\".\u002Fasserts\u002FTeaser_top.png\" alt=\"PixWorld teaser\" width=\"100%\">\n  \u003C\u002Fa>\n\u003C\u002Fp>\n\n\u003Cp align=\"left\">\n\u003Cstrong>TL;DR\u003C\u002Fstrong>: \u003Cstrong>PixWorld is a single end-to-end pixel-space diffusion model that unifies 3D scene generation and reconstruction\u003C\u002Fstrong> — it supervises a pixel-aligned 3D Gaussian field directly through differentiable rendering, with no VAE or RAE, and adds a geometry perception loss for 3D structural consistency.\n\u003C\u002Fp>\n\n\u003C\u002Fdiv>\n\n## ✨ Contributions\n\n- **One unified model for generation *and* reconstruction.** A single two-stream diffusion transformer processes posed multi-view inputs as a **clean** subset (→ reconstruction) and a **noisy** subset (→ generation, optionally text-conditioned), decoding a pixel-aligned 3D Gaussian scene in **one forward pass** — no task-specific branches.\n- **Pixel-space supervision, no VAE\u002FRAE.** A **flow-matching loss is imposed directly on rendered multi-view images** via differentiable rendering, so optimization is aligned with 3D scene fidelity instead of an intermediate latent target — removing the frozen VAE\u002FRAE and its reconstruction ceiling.\n- **Geometry perception loss.** Rendered views are aligned with ground truth in the geometry-aware feature space of a **frozen 3D foundation model (π³ \u002F VGGT)**, injecting 3D structural supervision beyond 2D photometric and perceptual losses.\n- **Real-time inference.** After distillation, the **4-step** model (`PixWorld-480P-4steps`) generates a scene in **~0.6 s** — up to **~1000×** faster than diffusion-based world generators.\n\n## 🎬 Showcase\n\n> One unified model, three capabilities — each grid shows six explorable 3D Gaussian scenes.\n> ▶️ **Full-quality, playable videos on the [project page](https:\u002F\u002Fsensengao.github.io\u002FPixWorld\u002F).**\n\n### 🏗️ 3D Reconstruction\n\nhttps:\u002F\u002Fgithub.com\u002Fuser-attachments\u002Fassets\u002F837b6898-6fe2-4b66-ac80-d2331eee71fe\n\n### 🖼️ Image → 3D\n\nhttps:\u002F\u002Fgithub.com\u002Fuser-attachments\u002Fassets\u002Fb1e681f6-d3e3-4703-87fe-90a3d9f5c922\n\n### ✍️ Text → 3D\n\nhttps:\u002F\u002Fgithub.com\u002Fuser-attachments\u002Fassets\u002F5353d7bf-5de3-4a3c-9b56-c49ca3db87ff\n\n## ⚡ Inference Speed\n\nA **single** PixWorld model performs both 3D reconstruction and generation. After distillation, the **4-step** model (`PixWorld-480P-4steps`) generates a scene in **~0.6 s** — up to **~1000×** faster than diffusion-based world generators (FantasyWorld 1041×, Gen3C 445×, Gen3R 148×, FlashWorld 5×).\n\n\u003Cp align=\"center\">\n  \u003Cimg src=\".\u002Fasserts\u002FSpeed.png\" alt=\"PixWorld inference speed comparison\" width=\"78%\">\n\u003C\u002Fp>\n\n## 🗓️ Release Plan\n\nWe plan to release the following **in a short time**:\n\n- [ ] 🧹 **Cleaned RealEstate10K \u002F DL3DV \u002F ACID datasets**\n- [ ] ⚡ **`PixWorld-480P-4steps` distilled model** — the 4-step distilled weights + inference code.\n\n## 🎓 Citation\n\nIf you find this repository useful, please consider citing PixWorld:\n\n```bibtex\n@misc{gao2026pixworld,\n      title={PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space},\n      author={Sensen Gao and Zhaoqing Wang and Qihang Cao and Dongdong Yu and Changhu Wang and Jia-Wang Bian},\n      year={2026},\n      eprint={2607.05373},\n      archivePrefix={arXiv},\n      primaryClass={cs.CV},\n      url={https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.05373},\n}\n```\n","PixWorld 是一个统一3D场景生成与重建的端到端像素空间扩散模型。它通过可微分渲染直接监督像素对齐的3D高斯场，摒弃VAE\u002FRAE等中间表示；采用双流扩散Transformer架构，单次前向即可处理多视角输入，支持无噪（重建）与加噪（生成，可文本引导）两种模式，并引入基于冻结3D基础模型（π³\u002FVGGT）的几何感知损失以增强结构一致性。模型经蒸馏后仅需4步采样，推理速度达约0.6秒\u002F场景，适合需高效、像素级保真度的3D内容创作、NeRF替代方案及多视角三维重建任务。","2026-07-11 02:30:47","CREATED_QUERY"]