[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-94353":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":9,"htmlUrl":10,"language":11,"languages":10,"totalLinesOfCode":10,"stars":12,"forks":13,"watchers":14,"openIssues":15,"contributorsCount":15,"subscribersCount":15,"size":15,"stars1d":15,"stars7d":15,"stars30d":16,"stars90d":15,"forks30d":15,"starsTrendScore":15,"compositeScore":17,"rankGlobal":10,"rankLanguage":10,"license":18,"archived":19,"fork":19,"defaultBranch":20,"hasWiki":19,"hasPages":19,"topics":21,"createdAt":10,"pushedAt":10,"updatedAt":22,"readmeContent":23,"aiSummary":24,"trendingCount":15,"starSnapshotCount":15,"syncStatus":25,"lastSyncTime":26,"discoverSource":27},94353,"RealReplicaBench","Accio-org\u002FRealReplicaBench","Accio-org","RealReplicaBench: Benchmarking Long-Horizon Agents in High-Fidelity, Stateful, and Reproducible Replicas of Real Online Services","https:\u002F\u002Frealreplicabench.site.accio.ai\u002F",null,"HTML",1040,79,1035,0,5,54.21,"Apache License 2.0",false,"main",[],"2026-08-24 04:01:22","\u003Cp align=\"center\">\n  \u003Cimg src=\"docs\u002Fassets\u002Frealreplicabench-banner.svg\" width=\"100%\" alt=\"RealReplicaBench — a stateful agent benchmark for real-world commerce workflows\">\n\u003C\u002Fp>\n\n\u003Cp align=\"center\">\n  \u003Ca href=\"https:\u002F\u002Fwww.accio.com\u002F\">\n    \u003Cimg src=\"docs\u002Fassets\u002Faccio-logo.svg\" height=\"30\" alt=\"Accio\">\n  \u003C\u002Fa>\n\u003C\u002Fp>\n\n\u003Cp align=\"center\">\n  \u003Cstrong>Developed and maintained by the Accio team at Alibaba International.\u003C\u002Fstrong>\n\u003C\u002Fp>\n\n\u003Cp align=\"center\">\n  \u003Ca href=\"#reproducibility-contract\">\u003Cimg alt=\"Release v1.3.1\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Frelease-v1.3.1-111827\">\u003C\u002Fa>\n  \u003Ca href=\"#task-and-run-layout\">\u003Cimg alt=\"107 tasks\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Ftasks-107-10b981\">\u003C\u002Fa>\n  \u003Ca href=\"#quick-start\">\u003Cimg alt=\"Python 3.11 or newer\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fpython-%E2%89%A53.11-00b2ff\">\u003C\u002Fa>\n  \u003Ca href=\"#quick-start\">\u003Cimg alt=\"OpenClaw harness\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fharness-OpenClaw-059669\">\u003C\u002Fa>\n  \u003Ca href=\"#reference-results\">\u003Cimg alt=\"OpenClaw and Accio reference results\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fresults-OpenClaw%20%2B%20Accio-047857\">\u003C\u002Fa>\n\u003C\u002Fp>\n\n\u003Cp align=\"center\">\n  \u003Ca href=\"#overview\">Overview\u003C\u002Fa> ·\n  \u003Ca href=\"https:\u002F\u002Frealreplicabench.site.accio.ai\u002F\">Live leaderboard\u003C\u002Fa> ·\n  \u003Ca href=\"https:\u002F\u002Frealreplicabench-mock-showcase.site.accio.ai\u002F\">Mock showcase\u003C\u002Fa> ·\n  \u003Ca href=\"#quick-start\">Quick start\u003C\u002Fa> ·\n  \u003Ca href=\"#reproducibility-contract\">Reproducibility\u003C\u002Fa> ·\n  \u003Ca href=\"#get-your-model-evaluated-or-work-with-us\">Contact\u003C\u002Fa>\n\u003C\u002Fp>\n\n\u003Cp align=\"center\">\n  \u003Ca href=\"#get-your-model-evaluated-or-work-with-us\">\n    \u003Cimg alt=\"Get your model evaluated on RealReplicaBench\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002F%E2%9C%89%20Get%20your%20model%20evaluated-047857?style=for-the-badge\">\n  \u003C\u002Fa>\n  &nbsp;\n  \u003Ca href=\"#get-your-model-evaluated-or-work-with-us\">\n    \u003Cimg alt=\"Collaborate with the Accio team on RealReplicaBench\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002F%F0%9F%A4%9D%20Collaborate-059669?style=for-the-badge\">\n  \u003C\u002Fa>\n\u003C\u002Fp>\n\n\u003Cp align=\"center\">\n  \u003Csub>We run models on request, including pre-release and internal builds —\n  and we are open to working together on the benchmark.\u003C\u002Fsub>\n\u003C\u002Fp>\n\n---\n\n## Overview\n\nRealReplicaBench evaluates whether an agent can complete long-horizon business\nworkflows, not just answer questions about them. Tasks cover browser operations,\nnative-style CLI tools, API\u002FMCP workflows, document and spreadsheet production,\npublic-web research, supplier analysis, product publishing, logistics, and\ncommerce operations. Every task runs in a fresh container and is graded by its\nown deterministic or LLM-assisted verifier.\n\n- **107 tasks:** 53 CLI, 28 browser, 16 file, and 10 API\u002FMCP tasks.\n- **Three capability slices:** 65 text-only, 20 browser-text-capable, and 22\n  vision-required tasks.\n- **Stateful evaluation:** local mock services model SaaS, commerce, messaging,\n  document, and operational systems without requiring production accounts.\n- **Auditable outputs:** each run preserves the resolved configuration,\n  trajectory, verifier result, artifacts, logs, and container metadata.\n\n\u003Cp align=\"center\">\n  \u003Cimg src=\"docs\u002Fassets\u002Fbenchmark-overview.svg\" width=\"100%\" alt=\"RealReplicaBench evaluation pipeline from business request to verified state change\">\n\u003C\u002Fp>\n\n### Real task surfaces\n\nThe suite uses reproducible local replicas of commerce and business software,\nso agents must operate interfaces and change state.\n\n\u003Ctable>\n  \u003Ctr>\n    \u003Ctd width=\"33%\">\u003Cimg src=\"docs\u002Fassets\u002Fscreenshots\u002Falibaba-publish-form.jpg\" alt=\"Product publishing workflow\">\u003C\u002Ftd>\n    \u003Ctd width=\"33%\">\u003Cimg src=\"docs\u002Fassets\u002Fscreenshots\u002Ffreightos-booking-search.jpg\" alt=\"Freight booking workflow\">\u003C\u002Ftd>\n    \u003Ctd width=\"33%\">\u003Cimg src=\"docs\u002Fassets\u002Fscreenshots\u002Fshopify-admin-theme-customize.jpg\" alt=\"Storefront theme customization workflow\">\u003C\u002Ftd>\n  \u003C\u002Ftr>\n  \u003Ctr>\n    \u003Ctd align=\"center\">\u003Cstrong>Product publishing\u003C\u002Fstrong>\u003Cbr>\u003Csub>Structured catalog and listing operations\u003C\u002Fsub>\u003C\u002Ftd>\n    \u003Ctd align=\"center\">\u003Cstrong>Freight booking\u003C\u002Fstrong>\u003Cbr>\u003Csub>Multi-step logistics workflows\u003C\u002Fsub>\u003C\u002Ftd>\n    \u003Ctd align=\"center\">\u003Cstrong>Storefront operations\u003C\u002Fstrong>\u003Cbr>\u003Csub>Visual configuration and stateful editing\u003C\u002Fsub>\u003C\u002Ftd>\n  \u003C\u002Ftr>\n\u003C\u002Ftable>\n\n\u003Cp align=\"center\">\n  \u003Ca href=\"https:\u002F\u002Frealreplicabench-mock-showcase.site.accio.ai\u002F\">\n    \u003Cimg alt=\"Explore the RealReplicaBench UI Mock Showcase\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FExplore-UI%20Mock%20Showcase-059669?style=for-the-badge\">\n  \u003C\u002Fa>\n\u003C\u002Fp>\n\n\u003Cp align=\"center\">\n  \u003Csub>Browse 104 rendered pages across eight UI mock services. The showcase is\n  a static visual tour; state-changing interactions run inside the benchmark\n  runtime.\u003C\u002Fsub>\n\u003C\u002Fp>\n\n## Reference results\n\nResults are aligned by `task_id` over the complete 107-task collection. The\ntables below are per harness — twelve model families on OpenClaw, thirteen on\nAccio — and the twelve present in both are the ones that compare directly.\nThe published scores were produced through Accio-managed evaluation endpoints\nwith `gemini-3.1-pro-preview` as the judge; the public path in this repository\nuses bring-your-own credentials.\n\nThe [live leaderboard](https:\u002F\u002Frealreplicabench.site.accio.ai\u002F) is the\nsource of record; the tables below are a snapshot.\n\n\u003Cp align=\"center\">\n  \u003Cimg src=\"docs\u002Fassets\u002Freference-leaderboard.svg\" width=\"100%\" alt=\"RealReplicaBench Leaderboard comparing OpenClaw and Accio\">\n\u003C\u002Fp>\n\n### Detailed evaluation statistics\n\nPass and capacity use the same verifier semantics across harnesses. Steps,\ntime, and tokens are descriptive telemetry: tool granularity, runtime\nscheduling, and provider usage accounting differ, so these values are not\nnormalized efficiency scores.\n\n🥇🥈🥉 mark the top three within each harness. The bar in the Pass column is\ndrawn on a fixed 0–100% scale, not normalized to the leader, so bar lengths\nare directly comparable between the two tables.\n\n#### OpenClaw\n\n| Model | Pass | Avg. capacity | Avg. steps | Avg. time | Avg. tokens |\n|---|:--|---:|---:|---:|---:|\n| 🥇 Claude Opus 5 | `███████████░░░░░░░░░` 60\u002F107 (56.1%) | 0.905 | 47.7 | 12.7 min | 3.47M |\n| 🥈 Claude Opus 4.8 | `██████████░░░░░░░░░░` 55\u002F107 (51.4%) | 0.860 | 47.6 | 16.4 min | 4.05M |\n| 🥉 GPT-5.6 Sol | `██████████░░░░░░░░░░` 53\u002F107 (49.5%) | 0.855 | 28.6 | 14.4 min | 2.09M |\n| GPT-5.5 | `██████████░░░░░░░░░░` 51\u002F107 (47.7%) | 0.835 | 37.1 | 12.7 min | 2.85M |\n| Claude Opus 4.7 | `█████████░░░░░░░░░░░` 49\u002F107 (45.8%) | 0.871 | 47.4 | 14.3 min | 4.10M |\n| Qwen 3.8 Max Preview | `█████████░░░░░░░░░░░` 48\u002F107 (44.9%) | 0.822 | 40.6 | 18.9 min | 2.13M |\n| Gemini 3.6 Flash | `█████████░░░░░░░░░░░` 48\u002F107 (44.9%) | 0.867 | 46.3 | 13.5 min | 3.28M |\n| DeepSeek V4 Flash | `█████████░░░░░░░░░░░` 46\u002F107 (43.0%) | 0.827 | 137.8 | 19.2 min | 11.04M |\n| GLM 5.2 | `████████░░░░░░░░░░░░` 42\u002F107 (39.3%) | 0.814 | 56.9 | 14.8 min | 3.12M |\n| Gemini 3.5 Flash | `███████░░░░░░░░░░░░░` 39\u002F107 (36.4%) | 0.798 | 63.9 | 17.9 min | 5.54M |\n| GPT-5.6 Luna | `███████░░░░░░░░░░░░░` 36\u002F107 (33.6%) | 0.797 | 27.5 | 12.2 min | 1.81M |\n| Gemini 3 Flash | `██████░░░░░░░░░░░░░░` 31\u002F107 (29.0%) | 0.744 | 45.1 | 16.1 min | 3.09M |\n\n#### Accio\n\n| Model | Pass | Avg. capacity | Avg. steps | Avg. time | Avg. tokens |\n|---|:--|---:|---:|---:|---:|\n| 🥇 Claude Opus 5 | `████████████░░░░░░░░` 66\u002F107 (61.7%) | 0.861 | 63.2 | 10.1 min | 3.69M |\n| 🥈 Claude Opus 4.8 | `███████████░░░░░░░░░` 59\u002F107 (55.1%) | 0.886 | 67.4 | 11.6 min | 4.82M |\n| 🥉 Claude Opus 4.7 | `██████████░░░░░░░░░░` 56\u002F107 (52.3%) | 0.878 | 61.5 | 6.4 min | 4.32M |\n| GPT-5.6 Sol | `██████████░░░░░░░░░░` 55\u002F107 (51.4%) | 0.873 | 53.0 | 5.5 min | 1.85M |\n| Qwen 3.8 Max | `██████████░░░░░░░░░░` 52\u002F107 (48.6%) | 0.826 | 67.8 | 15.6 min | 2.93M |\n| Gemini 3.6 Flash | `█████████░░░░░░░░░░░` 50\u002F107 (46.7%) | 0.815 | 47.7 | 4.6 min | 2.62M |\n| GLM 5.2 | `█████████░░░░░░░░░░░` 50\u002F107 (46.7%) | 0.787 | 81.0 | 10.8 min | 3.62M |\n| DeepSeek V4 Flash | `█████████░░░░░░░░░░░` 50\u002F107 (46.7%) | 0.838 | 84.0 | 10.0 min | 5.35M |\n| Qwen 3.8 Max Preview | `█████████░░░░░░░░░░░` 49\u002F107 (45.8%) | 0.856 | 69.8 | 12.7 min | 2.51M |\n| GPT-5.5 | `█████████░░░░░░░░░░░` 48\u002F107 (44.9%) | 0.864 | 45.3 | 4.5 min | 1.44M |\n| GPT-5.6 Luna | `█████████░░░░░░░░░░░` 48\u002F107 (44.9%) | 0.809 | 66.0 | 5.7 min | 2.49M |\n| Gemini 3.5 Flash | `█████████░░░░░░░░░░░` 46\u002F107 (43.0%) | 0.821 | 91.2 | 9.0 min | 4.80M |\n| Gemini 3 Flash | `██████░░░░░░░░░░░░░░` 31\u002F107 (29.0%) | 0.769 | 46.0 | 4.5 min | 2.48M |\n\nThe raw task-level result bundles are not stored in Git and do not yet have\npublic immutable URLs or checksums. Until they do, the published board is an\naudited aggregate keyed by public result IDs, not a standalone reproduction\npackage.\n\n### Get your model evaluated, or work with us\n\n> [!TIP]\n> **We run models on request**, including pre-release and internal builds, and\n> can evaluate privately against your own checkpoint before you ship it.\n>\n> **We are also open to collaboration** — new task domains, mock environments,\n> harness work, or joint evaluation. Tell us what you have in mind.\n>\n> [![Email Yukun Lian](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FYukun%20Lian-lianyukun.lyk%40alibaba--inc.com-059669?style=for-the-badge)](mailto:lianyukun.lyk@alibaba-inc.com)\n> [![Email Sicong Xie](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FSicong%20Xie-sicong.xsc%40alibaba--inc.com-047857?style=for-the-badge)](mailto:sicong.xsc@alibaba-inc.com)\n>\n> Prefer to copy rather than click: `lianyukun.lyk@alibaba-inc.com` ·\n> `sicong.xsc@alibaba-inc.com`\n\n### Metrics\n\n| Metric | Definition |\n|---|---|\n| Pass | A task passes only when every required verifier check passes; the rate is passes over the 107 aligned tasks. |\n| Avg. capacity | Macro mean of each task's `checks_passed \u002F checks_total`; this preserves partial task completion but is not a weighted official score. |\n| Avg. steps | Mean trajectory tool-call count over the displayed task attempts. |\n| Avg. time | Mean task wall-clock duration, using summary duration or audited manifest timestamps when the summary duration is zero. |\n| Avg. tokens | Mean total model tokens per task after normalizing provider-specific usage fields; cached tokens are included when reported. |\n\n## Quick start\n\n### Requirements\n\n- Docker with Linux container support (`linux\u002Famd64`; Apple Silicon hosts can\n  use emulation).\n- Python 3.11 or newer.\n- A model API key and an LLM-judge API key.\n\n### Install\n\n```bash\npython3 -m venv .venv\nsource .venv\u002Fbin\u002Factivate\npython -m pip install -e .\nreal-replica-bench list\n```\n\n### Pull the pinned OpenClaw runtime\n\nThe human-readable tag is mutable, so evaluation commands pin the current\nrelease digest:\n\n```bash\ndocker pull --platform linux\u002Famd64 \\\n  acciolyk\u002Faccio_bench@sha256:1e9cf5c72a56794175b7d06ece036b92e296e6b7e9e9a7fa244026f6acea3859\n```\n\nThe image contains OpenClaw `2026.5.22`, the browser stack, and the isolated\ndomain mock suite.\n\n### Run one task\n\nThis example uses Gemini's native `generateContent` path and the public Google\nAPI:\n\n```bash\nexport GEMINI_API_KEY=\"...\"\n\nreal-replica-bench run api-amazon-margin-floor-audit \\\n  --harness openclaw \\\n  --image acciolyk\u002Faccio_bench@sha256:1e9cf5c72a56794175b7d06ece036b92e296e6b7e9e9a7fa244026f6acea3859 \\\n  --platform linux\u002Famd64 \\\n  --openclaw-model google\u002Fgemini-3.5-flash \\\n  --openclaw-image-model google\u002Fgemini-3.5-flash \\\n  --openclaw-models-config configs\u002Frealreplicabench_native_google_direct_models.json \\\n  --llm-judge-provider gemini \\\n  --llm-judge-model gemini-3.1-pro-preview \\\n  --run-id realreplicabench-smoke\n```\n\n### Run a collection\n\n```bash\nreal-replica-bench run \\\n  --config configs\u002Frealreplicabench_openclaw_native_google_direct.yaml \\\n  --run-id \"realreplicabench-openclaw-$(date +%Y%m%d-%H%M%S)\"\n```\n\nUse `--limit 1` for a batch-path smoke test. The full suite can be partitioned\nwith the `*_text_only`, `*_browser_textcapable`, and `*_vision` collection\nfiles under `datasets_domain_v1\u002F`.\n\n### Provider routes\n\nEvery route is one config file in `configs\u002F`, all named\n`realreplicabench_openclaw\u003Csuffix>.yaml`. The tables list the suffix.\n\n**Managed routes** — a provider's own API, billed to that provider's key.\n\n| Route | Suffix | Wire protocol | Credentials |\n|---|---|---|---|\n| Native Gemini | `_native_google_direct` | Gemini `generateContent` | `GEMINI_API_KEY` |\n| Native Qwen \u002F DashScope | `_qwen37plus_native` | DashScope OpenAI-compatible | `DASHSCOPE_API_KEY` |\n| OpenRouter | *(none)* | OpenRouter chat, bundled shim | `OPENROUTER_API_KEY` |\n| Qwen through OpenRouter | `_qwen37plus_openrouter` | OpenRouter chat, bundled shim | `OPENROUTER_API_KEY` |\n| Custom native Gemini | `_native_google` | Gemini `generateContent` | Provider-specific |\n\n**Bring your own endpoint** — point OpenClaw at any base URL you control that\nspeaks one of these four wire formats, and evaluate a self-hosted or\npre-release model.\n\n| Wire format | Suffix | Credentials |\n|---|---|---|\n| OpenAI `\u002Fv1\u002Fchat\u002Fcompletions` | `_openai_chat` | `OPENAI_API_KEY`, or your endpoint's var |\n| OpenAI `\u002Fv1\u002Fresponses` | `_openai_responses` | `OPENAI_API_KEY`, or your endpoint's var |\n| Anthropic `\u002Fv1\u002Fmessages` | `_anthropic_messages` | `ANTHROPIC_API_KEY`, or your endpoint's var |\n| Gemini `generateContent` | `_custom_gemini` | `CUSTOM_GEMINI_BASE_URL` + `CUSTOM_GEMINI_API_KEY` |\n\nOverride the endpoint with `baseUrl` in the models JSON,\n`--openclaw-provider-base-url` (`--openclaw-base-url` for OpenRouter), or\n`--openclaw-api` to skip the preset entirely — see\n[`docs\u002Fopenclaw-byo-endpoint.md`](docs\u002Fopenclaw-byo-endpoint.md).\n\nThe judge is configured independently of the agent, on Gemini `generateContent`\nor the OpenAI Responses API. Six tasks include LLM-assisted checks; keep the\njudge on `gemini-3.1-pro-preview` unless you report a different one.\n\nSupply credentials through environment variables: the batch runner redacts them\nfrom `run.yaml` and fails on unresolved `${...}` placeholders before a container\nstarts. Evaluated models run with shell access to their container — see\n[`SECURITY.md`](SECURITY.md) for the key-handling rules that implies.\n\n### API validation boundary\n\nEvery route above — the Gemini, Qwen, and OpenRouter agents and both Judges,\nincluding reasoning through the bundled shim and custom upstream base URLs —\nhas been exercised against local protocol recorders, without real credentials\nor billable calls, under the exact request\u002Fresponse contracts covered by\n`tests\u002Ftest_public_api.py`.\n\nThis proves request construction and response parsing, not provider-side model\nentitlement, quota, or billing. Before a full run, use `--limit 1` with your own\nkeys and record the provider\u002Fmodel snapshot in the run metadata.\n\nDeeper reference:\n[`docs\u002Fopenclaw-runtime-image.md`](docs\u002Fopenclaw-runtime-image.md) for the\nruntime image's identity, pin, and customization boundary;\n[`docs\u002Fopenclaw-native-gemini.md`](docs\u002Fopenclaw-native-gemini.md) and\n[`docs\u002Fopenclaw-native-qwen.md`](docs\u002Fopenclaw-native-qwen.md) for the native\nprovider routes.\n\n## Reproducibility contract\n\nComparable runs pin these four:\n\n| Component | v1.3.1 pin |\n|---|---|\n| Task set | `realreplicabench_domain_v1_all` — 107 task IDs |\n| Task definitions | This repository release, including task workspaces and graders |\n| Harness | OpenClaw runner in this repository |\n| Runtime | `acciolyk\u002Faccio_bench@sha256:1e9cf5c72a56794175b7d06ece036b92e296e6b7e9e9a7fa244026f6acea3859` |\n\nReport the rest: provider, exact model and judge identifiers, endpoint class,\nreasoning configuration, task count, retry policy, and aggregation rule.\nCompare results only within one benchmark version — a release can change what a\ntask accepts — and never by displayed model name alone: routing, model\nsnapshots, prompt adapters, retry policies, and judge endpoints all change\noutcomes.\n\n## Task and run layout\n\n```text\ndatasets_domain_v1\u002F\n├── realreplicabench_domain_v1_{all,text_only,browser_textcapable,vision}.collection.json\n└── \u003Cinterface>\u002F\u003Cplatform>\u002F\u003Ctask>\u002F\n    ├── task.toml   task.md   workspace\u002F          agent-visible\n    └── grader\u002F     services\u002F private\u002F rubric.json\n```\n\nOnly `task.md` and `workspace\u002F` are staged into the agent-visible task tree;\ngraders, rubrics, private seeds, service launchers, and mock source stay\noutside it, and final artifacts go to `\u002Ftask\u002Foutputs\u002F`. After the agent exits,\nthe host-side verifier reads those outputs and the isolated mock state, writes\nthe reward record, archives logs and trajectories, removes the container, and\nleaves:\n\n```text\nruns\u002F\u003Crun_id>\u002F\n├── run.yaml   summary.json   summary.md   report.html\n└── tasks\u002F\u003Cindex>-\u003Ctask_id>\u002F\n    └── manifest.json  agent\u002F  verifier\u002F  workspace\u002Foutputs\u002F  screenshots\u002F  container\u002F\n```\n\n## Contributing\n\n**We are asking for your mock environments.** A benchmark with a fixed task set\ndecays: models saturate it and its answers drift into training data. Each new\nreplica service — a real service's API semantics, state transitions, and above\nall its rejections, running offline and scored deterministically — is a family\nof tasks no model has been trained on. The fourteen shipping today are\nregistered in `real_replica_bench\u002Fmock_services\u002Fregistry.py`.\n\n[`CONTRIBUTING.md`](CONTRIBUTING.md) states the bar a new mock has to clear, and\nthe rules for task fixes, graders, and harness changes. A merged mock reaches\nthe published benchmark when maintainers next rebake the runtime image.\n\nThe Accio team at Alibaba International built the harness, the mock services,\nand the v1 task suite; your pull request adds you to\n[`CONTRIBUTORS.md`](CONTRIBUTORS.md) alongside the mock itself.\n\nReport vulnerabilities privately per [`SECURITY.md`](SECURITY.md); third-party\nprovenance is inventoried in\n[`THIRD_PARTY_NOTICES.md`](THIRD_PARTY_NOTICES.md).\n\n## Citation\n\nCitation metadata is available in [`CITATION.cff`](CITATION.cff). Cite\n**RealReplicaBench (Accio)** together with release `v1.3.1` and the exact Git\ncommit used for evaluation. Until the accompanying paper is published, cite\nthe repository directly:\n\n```bibtex\n@misc{Lian2026RealReplicaBench,\n    author={Yukun Lian and Lei Wei and Sicong Xie and Guannan Zhang and Kesu\n            Wang and Hongyu Li and Chenhao Jiang and Lanbo Lin and Tianyuan\n            Yang and Xiaoyu Guo and Li Cai and Jialong Zhu},\n    title={RealReplicaBench: A Stateful Agent Benchmark for Long-Horizon Commerce and Business Workflows},\n    note={GitHub repository, v1.3.1},\n    howpublished={\\url{https:\u002F\u002Fgithub.com\u002FAccio-org\u002FRealReplicaBench}},\n    year={2026}\n}\n```\n\n## License\n\nRealReplicaBench is **open source**. It ships under two licenses, split the\nsame way as the repository itself:\n\n| Scope | License | File |\n|---|---|---|\n| Harness, Python package, mock-service code, scripts, and configs | Apache License 2.0 | [`LICENSE`](LICENSE) |\n| Task suite under `datasets_domain_v1\u002F` (task definitions, workspaces, graders, rubrics) | Creative Commons Attribution 4.0 International (CC BY 4.0) | [`LICENSE-DATA`](LICENSE-DATA) |\n\nCommercial use is allowed. Keep the license and attribution notices, state\nsignificant changes, and credit the benchmark as described under\n[Citation](#citation). Neither license grants trademark rights — \"Accio\" and\n\"RealReplicaBench\" identify this benchmark, not a fork of it.\n\n> [!IMPORTANT]\n> These terms cover **Accio's own contributions only**. The repository also\n> contains mirrored stylesheets, webfonts, icons, and recorded API responses\n> whose rights their owners retain. Every one is inventoried by owner and path\n> in [`THIRD_PARTY_NOTICES.md`](THIRD_PARTY_NOTICES.md); read it before\n> redistributing.\n","RealReplicaBench 是一个面向长周期、状态化智能体的基准测试框架，用于评估AI代理在高保真、可复现的线上服务仿真环境中执行真实商业工作流的能力。其核心功能包括107个覆盖CLI、浏览器操作、文件处理与API\u002FMCP调用的多样化任务，支持文本、文本+浏览器、视觉三类能力切片；采用本地容器化沙箱与专用确定性\u002FLLM辅助验证器，保障状态一致性与结果可审计。适用于大模型智能体研发团队对自动化办公、电商运营、SaaS集成等复杂业务场景下的代理能力进行系统性评测与迭代优化。",2,"2026-08-07 02:30:05","CREATED_QUERY"]