[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-96171":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":9,"htmlUrl":10,"language":11,"languages":10,"totalLinesOfCode":10,"stars":12,"forks":13,"watchers":14,"openIssues":15,"contributorsCount":15,"subscribersCount":15,"size":15,"stars1d":15,"stars7d":15,"stars30d":16,"stars90d":15,"forks30d":15,"starsTrendScore":15,"compositeScore":17,"rankGlobal":10,"rankLanguage":10,"license":18,"archived":19,"fork":19,"defaultBranch":20,"hasWiki":19,"hasPages":21,"topics":22,"createdAt":10,"pushedAt":10,"updatedAt":27,"readmeContent":28,"aiSummary":29,"trendingCount":15,"starSnapshotCount":15,"syncStatus":30,"lastSyncTime":31,"discoverSource":32},96171,"Show-Harness","showlab\u002FShow-Harness","showlab","Just a VLM Agent Can Play Robots 🦾","https:\u002F\u002Fshowlab.github.io\u002FShow-Harness\u002F",null,"Python",258,9,1,0,113,3,"Apache License 2.0",false,"main",true,[23,24,25,26],"embodied-agent","foundation-models","harness","robotics","2026-09-21 02:04:30","\u003Cp align=\"right\">\n  \u003Cb>English\u003C\u002Fb> | \u003Ca href=\".\u002FREADME.zh-CN.md\">简体中文\u003C\u002Fa>\n\u003C\u002Fp>\n\n\u003Cdiv align=\"center\">\n\n\u003Cimg src=\"assets\u002Fshow-harness-logo.svg\" alt=\"Show-Harness\" width=\"380\">\n\n### Just a VLM Agent Can Play Robots\n\n\u003Cp align=\"center\">\n\u003Ca href=\"https:\u002F\u002Fchenanno.github.io\u002F\">Yanzhe Chen\u003C\u002Fa>\u003Csup>*\u003C\u002Fsup> ·\n\u003Ca href=\"https:\u002F\u002Fwww.baizechen.site\u002F\">Zechen Bai\u003C\u002Fa>\u003Csup>*\u003C\u002Fsup> ·\n\u003Ca href=\"https:\u002F\u002Fcaozhijun.top\u002F\">Zhijun Cao\u003C\u002Fa>\u003Csup>*\u003C\u002Fsup> ·\n\u003Ca href=\"https:\u002F\u002Fwenzhengzeng.github.io\u002F\">Wenzheng Zeng\u003C\u002Fa>\u003Csup>*\u003C\u002Fsup> ·\n\u003Ca href=\"https:\u002F\u002Fqhlin.me\u002F\">Kevin Qinghong Lin\u003C\u002Fa>\u003Cbr>\n\u003Ca href=\"https:\u002F\u002Flinyq17.github.io\u002F\">Yiqi Lin\u003C\u002Fa> ·\n\u003Ca href=\"https:\u002F\u002Fethanliang99.github.io\u002F\">Guoqiang Liang\u003C\u002Fa> ·\n\u003Ca href=\"https:\u002F\u002Fkevinskwk.github.io\u002F\">Kevin Yuchen Ma\u003C\u002Fa> ·\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FceilingFan456\u002F\">Qiming Huang\u003C\u002Fa> ·\n\u003Ca href=\"https:\u002F\u002Fsites.google.com\u002Fview\u002Fshowlab\">Mike Zheng Shou\u003C\u002Fa>\u003Csup>&dagger;\u003C\u002Fsup>\n\u003C\u002Fp>\n\n\u003Cp align=\"center\">\u003Csup>*\u003C\u002Fsup> equal contribution &nbsp;·&nbsp; \u003Csup>&dagger;\u003C\u002Fsup> corresponding author\u003C\u002Fp>\n\n**Show Lab @ National University of Singapore**\n\n\u003Cp align=\"center\">\n📄 \u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.10522\">Arxiv Paper\u003C\u002Fa> &nbsp;|&nbsp;\n🤗 \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002Fpapers\u002F2609.10522\">Daily Paper\u003C\u002Fa> &nbsp;|&nbsp;\n🦾 \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002Fshowlab\u002FShow-Harness-VLMs\">Models\u003C\u002Fa> &nbsp;|&nbsp;\n📊 \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002Fdatasets\u002Fshowlab\u002FShow-Harness-Data\">Dataset\u003C\u002Fa> &nbsp;|&nbsp;\n🌐 \u003Ca href=\"https:\u002F\u002Fshowlab.github.io\u002FShow-Harness\u002F\">Project Page\u003C\u002Fa> &nbsp;|&nbsp;\n💬 \u003Ca href=\"https:\u002F\u002Fx.com\u002FZechenBai\u002Fstatus\u002F2097879130356498603\">X (Twitter)\u003C\u002Fa>\n\u003C\u002Fp>\n\n\u003C\u002Fdiv>\n\n\u003C!-- A bare user-attachments URL on its own line is the only form GitHub renders as a video\n     player; relative paths inside \u003Cvideo> are never rewritten. The same reel is committed at\n     assets\u002Fshow-harness-demo.mp4 for offline readers and forks. -->\n\nhttps:\u002F\u002Fgithub.com\u002Fuser-attachments\u002Fassets\u002Fbd2d31db-f5c5-4554-85bb-2aa206876ac7\n\n---\n\n## 🔥 News\n\n- [x] `2026.09` Beyond Show-Harness, we release [Awesome Multimodal Embodied Agents](https:\u002F\u002Fgithub.com\u002Fshowlab\u002FAwesome-Multimodal-Embodied-Agent), our survey of the **Agent + Robot** landscape from computer-use to **robot-use**.\n- [x] `2026.09` Public release: the harness, GUMI collectors, the plugin suite, and the training pipeline.\n- [x] `2026.09` Six LoRA adapters on [🤗 Show-Harness-VLMs](https:\u002F\u002Fhuggingface.co\u002Fshowlab\u002FShow-Harness-VLMs) and the demonstration corpus on [🤗 Show-Harness-Data](https:\u002F\u002Fhuggingface.co\u002Fdatasets\u002Fshowlab\u002FShow-Harness-Data).\n\n---\n\n## 📑 Table of Contents\n\n- [🌟 Overview](#-overview)\n- [🚀 Quick Start](#-quick-start)\n  - [1. Environments](#1-environments)\n  - [2. Collect demonstrations with GUMI](#2-collect-demonstrations-with-gumi)\n  - [3. Run a real robot](#3-run-a-real-robot)\n  - [4. Repository layout](#4-repository-layout)\n- [🤖 Two modes, one interface](#-two-modes-one-interface)\n- [📦 Released checkpoints and data](#-released-checkpoints-and-data)\n- [🧩 Plugins](#-plugins)\n- [🙏 Acknowledgements](#-acknowledgements)\n- [📌 Citation](#-citation)\n\n---\n\n## 🌟 Overview\n\n\u003Cp align=\"center\">\n  \u003Cimg src=\"assets\u002Foverview.png\" alt=\"Show-Harness overview\" width=\"92%\">\n\u003C\u002Fp>\n\n**Show-Harness** is an *embodied harness*: a compact semantic interface that lets a vision-language model **\"play\" a robot**. The model reasons over discrete, incremental action units; embodiment-specific interpreters ground each unit into motion, deterministically — so the VLM stays directly responsible for every physical decision.\n\nThrough the same interface, a closed-source frontier VLM controls a robot **zero-shot**, and a small open model becomes a capable policy with **less than a few H200 GPU-hours** of fine-tuning.\n\n- 🤖 **Two modes, one interface** — a frontier VLM zero-shot, or a fine-tuned small VLM emitting one action token per step.\n- 🦾 **Embodiment-agnostic** — Franka, AgileX Piper (single and dual arm), ManiSkill, and Isaac Lab share one vocabulary and one prompt set.\n- 🎮 **GUMI** — demonstrate a task by playing the robot in a browser; no teleoperation hardware, no post-processing.\n- 🧩 **Ablation-grade plugins** — one directory, one boolean, and byte-identical to no plugin when disabled.\n\n---\n\n## 🚀 Quick Start\n\n### 1. Environments\n\nSeparate venvs, because their pins conflict. Start with `base`; add the rest only when\nyou need them. `bash scripts\u002Fsetup.sh` with no arguments prints which already exist.\n\n| | build it with | what it is for |\n| --- | --- | --- |\n| `.venv` | `bash scripts\u002Fsetup.sh base` | the harness: collect, run a robot, drive a served VLM |\n| `.venv-vllm` | `bash scripts\u002Fsetup.sh serve` | serving a VLM locally (`scripts\u002Fserve_vlm.sh`) |\n\n`bash scripts\u002Fsetup.sh base --real` adds the Franka\u002FPiper hardware layer (RealSense, ROS\nshims, teleop window). Zero-shot and sim work do not need it.\n\nServing is its own process, so the harness talks to any OpenAI-compatible endpoint — a\nhosted model, or a colleague's server — without `.venv-vllm` existing at all.\n\nTraining is self-contained under [train\u002F](train\u002F) and builds its own venvs against upstream LLaMA-Factory (`bash train\u002Fscripts\u002Fsetup_llamafactory.sh`); nothing in the sections above depends on it.\n\n### 2. Collect demonstrations with GUMI\n\n\u003C!-- A loop of the interface driving itself, small enough to play inline. The full\n     rollout is committed at assets\u002Fgumi-rollout.mp4; GitHub will not render a\n     \u003Cvideo> that points at a repository path, so the reel here is a GIF. -->\n\u003Cp align=\"center\">\n  \u003Ca href=\"assets\u002Fgumi-rollout.mp4\">\n    \u003Cimg src=\"assets\u002Fgumi-rollout.gif\" alt=\"GUMI — a GUI agent driving both arms through the action units\" width=\"92%\">\n  \u003C\u002Fa>\n\u003C\u002Fp>\n\nGUMI maps every action unit to a key or button, so a human — or a GUI-driving\nagent — demonstrates a task by playing the robot in the browser, and every step\nis recorded as a training-ready (observation, action) pair. It drives the real\nrigs below; a synthetic tabletop world (`--sim`) lets you try the interface\nbefore any hardware is set up:\n\n```bash\nbash scripts\u002Fsetup.sh base\n.venv\u002Fbin\u002Fpython gumi\u002Fcollect_rollouts_web.py data\u002Frollouts_demo --sim\n# open http:\u002F\u002Flocalhost:8600 and drive the gripper with WASD \u002F arrow keys\n```\n\nThe same servers run against the real Franka\u002FPiper rigs (drop `--sim`), and the\nsame key bindings power live human takeover during autonomous rollouts. See\n[gumi\u002FREADME.md](gumi\u002FREADME.md) for the keyboard UI, the dual-arm UI, and the\nagent operators.\n\n### 3. Run a real robot\n\n1. Copy `configs\u002Fsite\u002Ffranka.yaml.example` to `configs\u002Fsite\u002Ffranka.yaml` and\n   fill in your robot address and camera serials (Piper:\n   `site\u002Fpiper_arms.yaml.example`).\n2. Copy `configs\u002Fsecrets.env.example` to `configs\u002Fsecrets.env` and add an API\n   key for the backend you use (`GEMINI_API_KEY` by default), or serve a local\n   VLM with `scripts\u002Fserve_vlm.sh`.\n3. Calibrate the safety floor and begin pose for your table — the shipped\n   values are examples, and every autonomous run refuses to descend below the\n   calibrated floor.\n4. Preflight — checks the environment, your site config and calibration, the\n   VLM backend (live), and that the robot and cameras answer:\n   `python scripts\u002Fcheck_setup.py --robot-config configs\u002Frobot_franka.yaml`\n5. `python scripts\u002Frun_real.py --robot-config configs\u002Frobot_franka.yaml`\n\nThe full walkthroughs are in [docs\u002Ffranka.md](docs\u002Ffranka.md) and\n[docs\u002Fpiper.md](docs\u002Fpiper.md); simulators in\n[docs\u002Fsimulators.md](docs\u002Fsimulators.md).\n\n### 4. Repository layout\n\n| Path | What it is |\n| --- | --- |\n| `core\u002F` | The interaction loop: runners for both modes, config layering, logging, the shared action vocabulary, and the provider-agnostic VLM client + roles (`core\u002Fvlm\u002F`) |\n| `plugins\u002F` | Harness plugins — each mounts on one stage of the loop and is toggled from the `plugins:` config block ([plugins\u002FREADME.md](plugins\u002FREADME.md)) |\n| `interpreters\u002F` | Embodiment interpreters: Franka (impedance), AgileX Piper (joint streaming), ManiSkill \u002F Isaac-Lab sims |\n| `gumi\u002F` | GUMI: browser teleoperation + agent operators; every step is recorded as a ready (observation, action) training pair |\n| `configs\u002F` | Layered configs: shipped defaults + your site identity + optional overlays ([configs\u002FREADME.md](configs\u002FREADME.md)) |\n| `prompts\u002F` | Controller prompts (zero-shot) and the versioned prompt contracts of fine-tuned checkpoints |\n| `scripts\u002F` | Rig bring-up, calibration capture, serving, data collection |\n| `train\u002F` | The fine-tuning pipeline: data conversion, dataset registration, LoRA configs ([train\u002FREADME.md](train\u002FREADME.md)) |\n| `models\u002F` | Chat templates, downloaded adapters, the HuggingFace cache ([models\u002FREADME.md](models\u002FREADME.md)) |\n| `docs\u002F` | Per-rig runbooks and the fine-tuned mode guide |\n\n---\n\n## 🤖 Two modes, one interface\n\n**Zero-shot** — a frontier VLM operates the full plugin harness with no\nrobot-specific training (`scripts\u002Frun_real.py`, `scripts\u002Frun_real_dual.py`).\n\n**Fine-tuned** — a small VLM fine-tuned on GUMI demonstrations emits one\naction token per step, planner-free (`scripts\u002Frun_real_mvtoken.py`). The real-robot\nconfigs default to `vlm_backend: qwen3_5_2b` (the `qwen3_5_2b_showharness_ft`\nadapter); serve your own checkpoint\ninstead and select it with `vlm_backend: finetuned_local`. See\n[docs\u002Ffinetuned.md](docs\u002Ffinetuned.md), including the training contracts\nthat must not drift.\n\nSwitching embodiments changes only the interpreter and its\n`configs\u002Fprimitives_\u003Cembodiment>.yaml`; the model-facing vocabulary and prompts\nstay the same.\n\n---\n\n## 📦 Released checkpoints and data\n\nFive LoRA adapters trained on the real corpus, one per backbone, at\n[showlab\u002FShow-Harness-VLMs](https:\u002F\u002Fhuggingface.co\u002Fshowlab\u002FShow-Harness-VLMs):\n`qwen3_5_0_8b`, `qwen3_5_2b`, `qwen3_5_4b`, `qwen3_5_9b`, `gemma4_e4b`; plus\n`qwen3_5_2b_sim`, one simulation policy covering both simulators. The\ndemonstrations they were trained on are at\n[showlab\u002FShow-Harness-Data](https:\u002F\u002Fhuggingface.co\u002Fdatasets\u002Fshowlab\u002FShow-Harness-Data)\n(real Franka\u002FPiper rollouts plus RoboLab and ManiSkill).\n\nFetch an adapter with the base model it needs, serve it, drive the robot:\n\n```bash\n# 1) fetch the adapter together with the base model it needs\nADAPTER=qwen3_5_2b WITH_BASE=1 bash scripts\u002Fmodel\u002Fdownload_vlm_model.sh\n\n# 2) serve it — the script activates .venv-vllm itself\nMODEL=Qwen\u002FQwen3.5-2B \\\n  LORA=qwen3_5_2b_showharness_ft=models\u002FShow-Harness-VLMs\u002Fqwen3_5_2b \\\n  FAMILY=qwen3_5 bash scripts\u002Fserve_vlm.sh\n\n# 3) drive the robot\npython scripts\u002Frun_real_mvtoken.py --robot-config configs\u002Frobot_franka_ft.yaml\n```\n\n`FAMILY` picks a jinja template from `models\u002Fchat_templates\u002F`. Training never reads one —\nLlamaFactory renders the conversation itself — so these exist only to make vLLM reproduce\nthat rendering at serve time. A base model's own template does not, and the mismatch fails\nsilently — see [models\u002FREADME.md](models\u002FREADME.md).\n\nTo fine-tune your own, [train\u002F](train\u002F) takes rollouts (yours or the released set) to a\nLoRA on any of the three supported families.\n\n---\n\n## 🧩 Plugins\n\nEach plugin mounts on one stage of the loop, is toggled by one boolean, and\nleaves the loop byte-identical when disabled.\n\n| Stage | Plugin (paper) | Code |\n| --- | --- | --- |\n| Perception | Multi-View Guidance | view-role prompt scaffolding + `core\u002Fprompting\u002Fwrist_marker.py` (`plugins\u002Fview_select` on the dual rig) |\n| Perception | Proprioception | `plugins\u002Fproprioception` |\n| Reasoning | Subtask Planning | `plugins\u002Fsubgoal` |\n| Reasoning | Situated Planning | `plugins\u002Fdeepplan` |\n| Reasoning | Action Chunking | `plugins\u002Faction_chunk` |\n| Reasoning | Adaptive Step | `plugins\u002Fvariable_step` |\n| Reasoning | Visual Prompt | `plugins\u002Faffordance` |\n| Action | Action History | `plugins\u002Fmem_text` |\n| Action | Failure Recovery | `plugins\u002Frecovery` (+ `plugins\u002Fauto_release` in the fine-tuned mode) |\n\n`plugins\u002FREADME.md` documents the contract for writing your own.\n\n---\n\n## 🙏 Acknowledgements\n\nShow-Harness builds on the following open-source work:\n\n- **Training** — [LLaMA-Factory](https:\u002F\u002Fgithub.com\u002Fhiyouga\u002FLLaMA-Factory)\n- **Serving** — [vLLM](https:\u002F\u002Fgithub.com\u002Fvllm-project\u002Fvllm)\n- **Franka control** — [Polymetis](https:\u002F\u002Ffacebookresearch.github.io\u002Ffairo\u002Fpolymetis\u002F)\n- **Simulation** — [ManiSkill](https:\u002F\u002Fgithub.com\u002Fhaosulab\u002FManiSkill), [Isaac Lab](https:\u002F\u002Fgithub.com\u002Fisaac-sim\u002FIsaacLab)\n- **Hardware SDK** — [AgileX Piper](https:\u002F\u002Fgithub.com\u002Fagilexrobotics)\n- **Open backbones** — Qwen3.5, Gemma 4, and InternVL3.5, which the released adapters are trained on\n\nThanks to all **[Show Lab @ NUS](https:\u002F\u002Fsites.google.com\u002Fview\u002Fshowlab)** members for their support.\n\n---\n\n## 📌 Citation\n\nIf you find Show-Harness useful, please cite:\n\n```bibtex\n@misc{chen2026showharnessjustvlmagent,\n      title={Show-Harness: Just a VLM Agent Can Play Robots}, \n      author={Yanzhe Chen and Zechen Bai and Zhijun Cao and Wenzheng Zeng and Kevin Qinghong Lin and Yiqi Lin and Guoqiang Liang and Kevin Yuchen Ma and Qiming Huang and Mike Zheng Shou},\n      year={2026},\n      eprint={2609.10522},\n      archivePrefix={arXiv},\n      primaryClass={cs.RO},\n      url={https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.10522}, \n}\n```\n\nIf you like the project, please give us a star ⭐ — it is how we hear that it is useful.\n\n\u003Ca href=\"https:\u002F\u002Fstar-history.com\u002F#showlab\u002FShow-Harness&Date\">\u003Cimg alt=\"Star History Chart\" src=\"https:\u002F\u002Fapi.star-history.com\u002Fsvg?repos=showlab\u002FShow-Harness&type=Date\">\u003C\u002Fa>\n\n","Show-Harness 是一个面向具身智能体的视觉语言模型（VLM）代理框架，旨在将多模态大模型能力接入真实机器人系统。其核心功能包括：统一接口支持仿真与真机双模式运行；提供 GUMI 工具链用于人类示范数据采集；集成可插拔的感知-规划-控制插件模块；并开源了适配多种机器人的 LoRA 微调模型及配套指令-动作对数据集。项目基于 Apache 2.0 协议，强调轻量部署与模块化扩展，适用于具身智能研究、机器人技能学习、VLM 驱动的物理交互验证等学术与原型开发场景。",2,"2026-09-11 02:30:11","CREATED_QUERY"]