[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-92453":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":9,"htmlUrl":10,"language":10,"languages":10,"totalLinesOfCode":10,"stars":11,"forks":12,"watchers":13,"openIssues":12,"contributorsCount":12,"subscribersCount":12,"size":12,"stars1d":12,"stars7d":12,"stars30d":12,"stars90d":12,"forks30d":12,"starsTrendScore":12,"compositeScore":14,"rankGlobal":10,"rankLanguage":10,"license":15,"archived":16,"fork":16,"defaultBranch":17,"hasWiki":18,"hasPages":18,"topics":19,"createdAt":10,"pushedAt":10,"updatedAt":27,"readmeContent":28,"aiSummary":29,"trendingCount":12,"starSnapshotCount":12,"syncStatus":30,"lastSyncTime":31,"discoverSource":32},92453,"Agentic-First-Enterprises","MihaiCiprianChezan\u002FAgentic-First-Enterprises","MihaiCiprianChezan","A reference operating model for organizations where any role can be filled by a human or an agent, every process is built for agents first, and every flow can be paused, adjusted by a human, and resumed.","",null,291,0,190,40,"Creative Commons Attribution 4.0 International",false,"main",true,[20,21,22,23,24,25,26],"ai-agents","ai-agents-automation","ai-agents-framework","enterprise-ai","enterprise-architecture","llm","multi-agent-systems","2026-07-22 04:02:06","---\nAuthor: Mihai-Ciprian Chezan\nVersion: 2.0.2\nDate: 2026-07-07\n---\n\n# The Agent-Native Enterprise\n\n**A reference operating model for organizations where any role can be filled by a human or an agent, every process is built for agents first, and every flow can be paused, adjusted by a human, and resumed.**\n\n![The-Agent-Native-Enterprise](.\u002Fimages\u002FThe-Agent-Native-Enterprise.webp)\n\nThis is a generic model. It is inspired by software delivery but is not specific to it; the same structure applies to any goal-pursuing organization (operations, services, research, manufacturing coordination, back-office).\n\nTwo properties define it and separate it from a conventional org chart with chatbots bolted on:\n\n1. **Role polymorphism** — a role is a contract, not a person. An agent, a human, or a deterministic program can implement the same contract interchangeably (how far interchangeability stretches at high throughput is stated precisely in invariant #2 and §7).\n2. **Universal interruptibility** — every autonomous flow is built with a handbrake. A human with AI skill can stop it, inspect its state, inject an adjustment, and resume — by design, on every flow, not as an exception.\n\n**Humans do not disappear from the top of this model.** They move to where human judgment and human accountability are irreplaceable. Above the agentic organization sits a human **Representation & Accountability layer** — the *Board* — the people who set the organization's purpose and boundaries, carry its legal and public responsibility, and answer for it in the human world. They are not removed; they are relocated *out* of the millisecond decision loop, where human latency would only be a bottleneck, and *into* the part of the structure only humans can hold. Crucially, this layer is a **pattern, not a headcount**: the Board may be a hundred people, ten, or a single founder. A one-person company and a large enterprise run the identical model — only the size of the human layer differs. Existing leadership structures map onto this layer intact; what changes is that they stop being the operational bottleneck. (See §3.)\n\n**The model is not limited to new organizations.** Its unit is a self-contained, sovereign **cell** — which can be an entire company *or* a single bounded zone running in parallel inside an existing organization, staffed by people who already work there. Many cells compose into a federation without changing the pattern. (See §16.)\n\n## In one paragraph\n\n*One model, one rule of thumb: humans hold the Offices and write the constitution; agents fill the Roles and run the work; system roles keep the machine healthy (Steward), efficient (Optimizer), and trustworthy across versions (Auditor); any flow can be stopped, corrected by hand, and resumed; and a human can become any Role at any moment — bound, like every agent, by the constitution they set.*\n\n*(Prefer narrative? A two-page overview and a twenty-line at-a-glance version live in [SUMMARY.md](SUMMARY.md).)*\n\n## Contents\n\n[§1 Design invariants](#1-design-invariants) · [§2 The role as an interface](#2-core-abstraction-the-role-as-an-interface) · [§3 The Representation & Accountability layer](#3-the-representation--accountability-layer-the-human-board) · [§4 Operating and system roles](#4-operating-and-system-roles) · [§5 The planes](#5-the-planes) · [§6 The Handbrake](#6-the-handbrake-control-plane-in-depth) · [§7 Agent-first, human-tolerant execution](#7-agent-first-human-tolerant-execution) · [§8 Authority and autonomy](#8-authority-and-autonomy-model) · [§9 The Steward](#9-the-steward-the-org-doctor-named-for-what-it-does) · [§10 The Optimizer](#10-the-optimizer-capability-to-task-matching) · [§11 The Auditor](#11-the-auditor-version-fitness-and-safety) · [§12 Escalation and human takeover](#12-escalation-and-human-takeover) · [§13 Reference topology](#13-reference-topology) · [§14 Failure modes, guardrails, and trust boundaries](#14-failure-modes-and-the-guardrails-that-contain-them) · [§15 Adoption sequence](#15-adoption-sequence-lean) · [§16 Cells, sovereignty, and federation](#16-the-cell-sovereignty-and-federation) · [§17 Constitutional mechanics](#17-constitutional-mechanics) · [§18 Worked examples](#18-worked-examples-end-to-end) · [§19 Related work](#19-related-work-positioning-and-references) — [Appendix A: the seven role contracts](#appendix-a--the-seven-role-contracts-normative-baseline) · [Appendix B: a minimal handoff vocabulary](#appendix-b--a-minimal-handoff-vocabulary-informative-proven-in-one-cell) · [Appendix C: conformance profiles and checklist](#appendix-c--conformance-profiles-and-checklist) · [Appendix D: glossary](#appendix-d--glossary-informative)\n\n## How to read this document\n\n**Conformance language.** Statements written with **must \u002F never \u002F is required** are binding: an implementation that lacks them does not conform. Statements with **should \u002F recommended** are strong defaults, deviated from only with recorded rationale. Statements with **may \u002F optional** are choices. The eleven invariants are cited as **INV-1…INV-11** (interchangeably, \"invariant #n\") and the §6 handbrake requirements as **HB-1…HB-4**. The one-paragraph summary, the contents, §13's diagram, §18's examples, §19, and Appendices B and D are informative. Appendix A is a **normative baseline** an adopting cell copies and may adapt (the adaptations are constitutional content). In Appendix C the **profile definitions are normative**; the checklist is a derived aid — where a checklist row and the body diverge, the body governs. Everything else is normative. Where the model deliberately leaves a value to each organization it says *declared in the constitution* — that is delegation, not an unspecified gap.\n\n**Maturity.** §1–§15 and §17–§18 are implemented at least once by the reference cell (see the note at the end). §16's federation layer — treaties, the supra-constitution — is design-stage: not yet exercised by any two-cell deployment, and honestly labeled speculative until it is.\n\n---\n\n## 1. Design invariants\n\nThese are the non-negotiable rules, cited throughout as INV-1…INV-11. Everything below is a consequence of them.\n\n1. **Depend on the contract, not the implementer.** No part of the system may assume a role is held by an agent or by a human. It depends only on the role's declared inputs, outputs, authority, and guarantees.\n2. **Agent-first, human-tolerant.** Default execution runs at agent speed. Every role can be *stopped, inspected, and corrected* by a human at any time; a role whose declared throughput is human-boundable can additionally be *run* by one, and a role that is not must declare suspension-plus-inspection as its human-takeover mode in its contract (§2, §7). Either way, the surrounding system must continue functioning at degraded speed without cascading failure.\n3. **Every flow has a handbrake.** Interruptibility is core architecture, not a feature. A flow that cannot be paused, inspected, adjusted, and resumed is non-compliant.\n4. **Side-effecting actions are made as safe to retry as the effect allows.** Where an effect is yours or reversible, it is idempotent — retry or resume never duplicates it. Where it is irreversible and owned by a non-idempotent outsider (you cannot un-send a message or un-ship a unit), the guarantee narrows to at-most-once *attempts* plus compensation where reversal exists. The outside world is never assumed idempotent; safety is engineered on the side you control.\n5. **State lives outside the actor.** Context, progress, and history are held in shared, durable storage — not inside the agent's transient memory or a human's head — so any implementer can take over a role mid-flow.\n6. **Authority is graduated and explicit.** Every action class has a declared autonomy level. Blast radius determines how much human gating it requires.\n7. **One abstraction per layer.** A layer never reaches across levels. The top layer does not know about individual tool calls; a worker does not know the global strategy. This is what keeps the system debuggable.\n8. **Add hierarchy only when complexity forces it.** Most designs overshoot by one tier. Start with the fewest layers that solve the problem.\n9. **Office ≠ Role.** A human holds an *Office* (accountability and representation in the human world); an agent fills a *Role* (operations in the agentic org). They are not one-to-one. The corollary is the **two-channel rule**: humans affect the agentic org through exactly two channels — the constitution (authoring and amending it, *and exercising the gate-powers it explicitly grants to humans*: an L1 approval, an L0 execution, a break-glass act — §8, §17), and impersonation of a Role — and nothing else. A gate-power is narrow, momentary, and enumerated; impersonation is open-ended holding of a Role's seat under the Role's authority. One invariant, two faces: the distinction says what a human *is* to the org; the corollary says how a human *reaches* it.\n10. **Governance is the compiled constitution — its rule-shaped part.** The Governance plane encodes the *projection* of the constitution that reduces to rules: authority ceilings, permissions, budget caps, required gates. The purposive core — whether the org still serves human interest — does not compile and stays in human Board review (§3). Every encoded constraint traces back to a written human mandate; agents never author their own constraints. The plane enforces the mechanical fraction; human judgment carries the substantive remainder.\n11. **A cell is sovereign at its boundary.** A cell is an organization in itself. Nothing outside it — a parent organization or a sibling cell included — may affect it except through its own constitution or an authorized Role, exactly as inside. The boundary obeys the same law as the interior.\n\n---\n\n## 2. Core abstraction: the role as an interface\n\nA **role** is a declared contract with this shape:\n\n| Field | Meaning |\n|---|---|\n| **Responsibility** | The single outcome this role owns. |\n| **Inputs** | What it consumes, and from which roles. |\n| **Outputs** | What it produces, and to which roles. |\n| **Authority scope** | What it may decide and act on alone; what it must escalate. |\n| **Acceptance criteria** | How \"done\" and \"correct\" are judged. |\n| **Escalation rule** | The conditions under which it must hand off to a human or a higher role. |\n| **Observability hooks** | The traces, costs, and signals it must emit. |\n\nAn **implementer** — agent, human, or deterministic program — satisfies the contract. The system binds to the contract. This makes \"a human jumps into the role\" a runtime substitution, not a redesign. It is the same principle as an interface with interchangeable implementations: the caller is unaffected by which one is running. The contract is given here as *fields and guarantees*, deliberately not as any particular file format or schema language: the model specifies what a role must declare, never how to write it down. That omission is intentional — it keeps the model independent of any toolchain and any era of tooling. A contract also declares its **human-takeover mode** — *run* (a human can hold the seat at its working throughput) or *suspend-and-inspect* (a human stops, inspects, and corrects it, but cannot run it live) — per INV-2 and §7. Appendix A instantiates this contract shape for all seven roles.\n\n**Who writes the acceptance criteria.** Criteria are authored by the role that *issues* the work — Direction sets them when it specifies a goal (turning demand into well-specified direction is that role's one job, §4.1), and each decomposition inherits or refines them downward. Two roles are barred from authoring them: the Executor (the producer cannot write its own bar) and the Verifier (the gate cannot either — it scores against criteria it did not set, which is what keeps the checker's independence real). Every criterion must be *checkable* — a statement the Verifier can score as met or unmet, not an aspiration — and *mechanically* checkable where possible (a test, a schema, a policy predicate): judgment-graded criteria are permitted, but they inherit the machine-judge reliability caveat of §4.4. When the Verifier finds a criterion untestable or ambiguous — or a goal whose framing contradicts the constitution or its own stated premises — it does not interpret silently: it returns the goal to Direction as an escalation, the same way it returns a failing output. Ambiguity is forced back to the role paid to resolve it, never absorbed by the role paid to judge.\n\n**Human impersonation on demand** is the default mode for substitution: every role runs as an agent unless and until a human assumes it for a specific need (a hard decision, a novel situation, a correction, an audit), then hands it back.\n\n---\n\n## 3. The Representation & Accountability layer (the human Board)\n\nThe agentic organization needs a human anchor in the human world — something an agent cannot be. An agent cannot hold legal accountability, sign a binding commitment, face a regulator, or stand as the responsible human face of the entity. Conventional organizations fuse this representation with day-to-day operational decisioning. This model **splits them**:\n\n- **Accountability and representation** → human, deliberately *outside the hot path*, operating at human-world cadence.\n- **Operational decisioning** → agentic, *inside the hot path*, at agent speed.\n\nThe reason is simply speed: a human permanently in the decision loop becomes the bottleneck the entire system must slow down to. So the humans sit above the loop and reach into it only on purpose.\n\n### The Board is a pattern, not a group\nThe Board may be a hundred people, ten, or one. A single-founder company runs a Board of one; a large enterprise runs a large one. The responsibilities below are identical at every size — only the headcount changes. Existing executive and governance structures fit here unchanged; they simply stop being the operational throughput limit.\n\n### Isn't the Board just the management layer, smuggled back in?\nA fair objection, given the model folds CEO, Product Owner, and Project Manager into one agentic Director: if agents absorb the coordination layer, why does a human layer reappear on top? Because the Board does the one thing an agent *structurally* cannot — it holds legal accountability and authors the constitution. \"Management\" in the sense the model removes is *operational coordination*: deciding who does what and when, reconciling the moment-to-moment — exactly what the Director and Orchestrator absorb. What the Board keeps is not coordination but *answerability and purpose-setting*: being the human the law and the public hold responsible, and writing the goals the system pursues. Those do not get more efficient at agent speed; they require a human because their value is accountability, not throughput. The Board is not the management layer returning — it is the residue left once coordination is automated away. \"Residue\" understates one thing, so it is stated here rather than discovered later: the Board also keeps the organization's *outward-facing* work — market sensemaking, capital, partnerships, the external relationships whose value is human trust — human-held for the same reason accountability is: their currency is trust, not throughput. Where relationship capital is the product, this function is large, and the Board should treat it as a duty, not a leftover.\n\n### What the Board does (levers only — never hands-on operation)\n1. **Authors the constitution** — the organization's purpose, values, goals, and behavioral boundaries. This is the source document the top agents operate under.\n2. **Maintains and amends it** — the constitution is living; the Board owns its changes.\n3. **Runs periodic human-interest alignment review** — checks that the agentic org still serves the human purpose it was built for, and has not drifted into doing something technically flawless but wrong. The test is concrete: divergence from the *written* purpose is drift, to be corrected; a deliberate, ratified change to that purpose is evolution, a constitutional amendment. The Board measures behavior against the text, not against a mood. The review's power is bounded by the independence of its evidence: a review fed solely by the cell's own telemetry is non-compliant — at least one channel the reviewed system cannot shape is required (Board-chosen raw-trace sampling from the tamper-evident event plane, direct stakeholder contact, or an external audit) — and cadence and sample size are declared in the constitution against the cell's action volume. The constitution also declares a ceiling on the interval (or action volume) between Board spot-reviews of Direction's goal-framing, so \"periodic\" bounds how long a mis-framed goal can run unexamined (§4.1).\n4. **Holds binding authority to mandate modification** — issued as a constitutional amendment or a formal change request, never as turn-by-turn meddling.\n5. **Carries legal and representational accountability** — answers for the entity in the human world.\n6. **Owns succession** — of Board members and of the standby humans who can competently impersonate each critical Role (§12). The model's safety valve is a human bench that decays without deliberate regeneration (§7); keeping it staffed and practiced is a Board duty, not an assumption.\n\n### How the Board itself decides\nThe model governs the Board the way it governs everything else: by writing it down. A Board of one needs no procedure; any larger Board must declare its *own* decision rules **in the constitution** — quorum, the threshold to ratify or amend, how internal disputes resolve, and what happens on deadlock. This is the model's self-referential closure: the constitution governs its own amendment process, so \"how the Board decides\" is never improvised at the moment it matters most. The model requires only that these rules exist and are written; it does not prescribe their content — a founder may keep sole authority, a council may require supermajority — that is the organization's choice.\n\n### Constitution → governance pipeline\nThe Board writes the **constitution** (human language, human intent). It is compiled into the machine-readable **Governance plane** (§5), which the top agents read and obey at runtime. This is the same shape as constitution → law → regulation: human principle becomes enforceable runtime constraint. It is also why invariant #10 holds — agents never write their own rules.\n\n### Office ≠ Role\nA human holds an **Office** — Board member, Chair, CEO, CFO — a human-world title carrying accountability and representation. An agent fills a **Role** — Director, Orchestrator, Executor, Verifier, Steward, Optimizer, Auditor — an operational seat in the agentic org. **They are distinct and not one-to-one.** \"CEO\" is an Office; \"Director\" is the Role that owns top-level operational direction. The human CEO does not *run* the org turn by turn; the Director agent does.\n\n### The impersonation-binding rule\nWhen a Board member impersonates a Role — say, steps into the Director seat through the handbrake — **they act as that Role**: they inherit the Director's authority scope and are bound by the same Governance plane as the agent would be. They do **not** carry their Office authority into the seat. A human cannot enter a Role and act outside the constitution \"because they are really the CEO.\" Doing something the governance forbids is not a keyboard action — it is a **constitutional amendment**, which is slow, deliberate, and audited. The handbrake lets humans operate *within* policy; only the Board, acting as the Board, changes policy. Two powers, two speeds, two separate audit trails.\n\n### Board review vs. Steward monitoring vs. Auditor rating (don't confuse them)\n- The **Steward** (§9) watches *operational\u002Fbehavioral* drift of a live instance: is the agent following the written rules and working correctly right now — continuous, technical, agent-speed.\n- The **Auditor** (§11) rates *versions* over accumulated activity: is this release better, worse, or more dangerous than the last — continuous, evaluative, agent-speed.\n- The **Board** watches *purpose* drift: are the rules still the right rules, is the org still serving human interest — periodic, judgment-based, human-speed.\n\nThe Steward catches \"the Director instance is malfunctioning.\" The Auditor catches \"Director v5 regressed against v4.\" The Board catches \"the Director is flawlessly executing the wrong mission.\" None can do the others' jobs.\n\n### The one Board failure mode to guard against\nThe Board must not slide into **shadow operation**. The moment it makes continuous turn-by-turn calls, it reintroduces the human bottleneck it exists to eliminate. Its influence flows through the constitution and through bounded Role-impersonation — not through meddling.\n\n---\n\n## 4. Operating and system roles\n\nRoles divide into two kinds. **Operating roles** form the value chain — they turn intent into delivered outcomes. **System roles** own no business outcome; they keep the machine itself healthy, efficient, and trustworthy. Each role lists a suggested human-readable holder name; the generic noun is the role, the name in parentheses is what a reader can picture a \"person\" being called. Roles are *logical contracts*, not necessarily separate systems: two roles may run on one underlying implementation with different permissions. What separates them is authority and the object they act on, not the process that runs them — with one caveat the trust boundaries (§14) make explicit: sharing an implementation separates *authority*, never failure modes, so a checking role should not share a single point of provenance with what it checks on high-blast-radius classes (§4.4, §10).\n\n### Operating roles\n\n#### 4.1 Direction *(holder: Director)* — consolidates CEO + Product Owner + Project Manager\nOwns **what and why**. Takes intent from stakeholders or clients, converts it into a prioritized backlog of goals with constraints and acceptance criteria, operating under the constitution. Highest operational decision authority. The three legacy roles collapse here on a stated bet: *to the extent* execution and coordination are competently agent-driven, the residual human-scarce input is clarity of intent — one job: turn demand into prioritized, well-specified direction. It is named as a bet because the measured evidence still cuts the other way for current-generation systems — inter-agent coordination failures remain a leading empirical failure class (§19: MAST) — which is exactly why Orchestration stays its own role, the Steward watches it hardest, and §14 keeps \"supervisor as single point of failure\" on the books. At large scale this role can concentrate too much; the answer is not to bloat the Director but to **split the cell** (§16) or, within INV-8, to tier Direction (portfolio-level over cell-level) — decompose only when the load actually demands it. Concentration is also a *framing* risk, not only a load risk: the Director authors both the goals and their acceptance criteria, so a mis-framed goal sails through every downstream gate while the Board's catch is periodic. Three cheap counters exist: the Verifier's premise-bounce (§2), the constitution-declared ceiling on the Board's spot-review interval (§3), and — for high-stakes goals — a **divergent-framing check**: a second Direction variant (or a human) frames the goal independently, and disagreement escalates (§12).\n\n#### 4.2 Orchestration *(holder: Orchestrator)*\nOwns **who does what, and when**. Decomposes each goal into work, routes it to Executors, sequences dependencies, handles exceptions, and decides whether to retry, escalate, or proceed. The supervisor layer. It does not do the work and does not set strategy.\n\n#### 4.3 Execution *(holder: Executor, a.k.a. Specialist)*\nOwns **how**. Specialist implementers that produce the actual work product. Narrow, deep, replaceable. An Executor knows its task and its tools, not the global plan.\n\n#### 4.4 Verification *(holder: Verifier, a.k.a. Reviewer)*\nOwns **is it correct and within policy**. Independently scores outputs against acceptance criteria, quality, safety, and conformance before they take effect. Runs as an evaluator loop against Execution: produce → score → revise. Independence from Execution is the point — the checker is not the producer. Role independence is not *statistical* independence, though: an Executor and a Verifier built on the same or a similar base model share blind spots, and machine judges measurably favor outputs of their own lineage (§19). For high-blast-radius classes the constitution should require the Verifier seat to be **implementation-independent** of Execution — a different model family, or best of all a deterministic checker. Where acceptance criteria are mechanically checkable, a deterministic program is the *preferred* Verifier: its independence is established by construction rather than monitored, it cannot be prompt-injected, and it is the cheapest implementer available — role polymorphism working as designed (§2), not an exception to it.\n\nThe gate's three outcomes are defined: **pass** — the output may take effect; **return** — revise against cited criteria, bounded by a revision limit declared in the constitution, after which the flow escalates per §12 (an unbounded produce→score→revise loop is optimization pressure against the gate — reward hacking needs no compromise, only iterations); **block** — a policy or boundary violation, a binding stop, logged with the cited clause. A per-criterion *unclear* score triggers §2's ambiguity rule: returned to Direction, never interpreted silently. Like any role, Verification may decompose at scale into distinct correctness, compliance, and risk checks — but only when complexity forces it (INV-8); by default it is one gate, and policy\u002Fcompliance is enforced cross-cuttingly by the Governance plane rather than bolted onto Verification.\n\n### System roles *(non-authoritative — they optimize the machine, not the business)*\n\n#### 4.5 Steward *(holder: Steward)* — the renamed \"org doctor\"\nOwns **are the role-holders themselves healthy and behaving normally**. Monitors, maintains, tunes, and repairs the other roles — especially the high-authority Direction and Orchestration agents. Full technical capability over them, **zero business-decision authority**. Detailed in §9.\n\n#### 4.6 Optimization *(holder: Optimizer)*\nOwns **is each task handled by the right-capability implementer**. Sits optionally between steps and decides which model\u002Fimplementer a given task requires, bounded by the task's risk and quality floor. Detailed in §10.\n\n#### 4.7 Audit *(holder: Auditor)*\nOwns **how versions of agents compare over time** — whether a given version is better, worse, or more dangerous than another, and which release is currently fit. Continuously rates agent versions from their accumulated activity; normally only monitors and reports. May **suspend** a dangerous or severely drifting version and escalate to humans, but cannot modify, dismiss, restart, or reinstate. Detailed in §11.\n\n| Kind | Role | Holder | Owns | Authority |\n|---|---|---|---|---|\n| Operating | Direction | Director | What & why | Highest (within constitution) |\n| Operating | Orchestration | Orchestrator | Who\u002Fwhen | Routing, retry, escalate |\n| Operating | Execution | Executor | How | Within-task only |\n| Operating | Verification | Verifier | Correctness\u002Fconformance | Gate: pass \u002F return \u002F block |\n| System | Steward | Steward | Health of live role-holders | Technical only, **no business decisions** |\n| System | Optimization | Optimizer | Capability-to-task fit | Selection only, **no business decisions** |\n| System | Audit | Auditor | Fitness & safety of versions | Monitor\u002Freport; may suspend + escalate; **no modify, dismiss, or reinstate** |\n\nEvery role: agent by default, human impersonation on demand — and a deterministic program wherever the contract is machine-checkable (§4.4).\n\n---\n\n## 5. The planes\n\nRoles run *inside* four cross-cutting planes. Planes are shared infrastructure every role depends on.\n\n- **Governance plane** — the machine-readable, runtime-enforced encoding of the Board's constitution (§3): authority limits, guardrails, constraints, and an append-only audit trail. Policy is data the agents read at runtime, not documentation humans read later. Nothing acts outside it.\n- **Memory \u002F context plane** — durable shared state: the current state of every goal, an append-only event history of decisions and actions, the artifacts produced, and a **version registry** that identifies every agent version and parallel variant and attributes activity to it (without which the Auditor cannot compare versions). A *version* here is the whole behavioral bundle — logic, prompts, weights, configuration — not merely code; and a human takeover or handbrake adjustment **opens a tracked variant** (a variant_of the incumbent version) for the duration of human control, closed on hand-back — never an untracked change — so the registry stays the source of truth, a version's scorecard measures the release rather than a human quietly rescuing it, and human interventions are separately attributable. Stated as fields and guarantees (like the role contract, §2 — never a schema language): a registry entry declares *identity* (role, version id — and, where the implementer is a hosted model, the pinned upstream model snapshot, or provider-side updates silently change the behavioral bundle mid-rating), *lineage* (what it derives from — predecessor or the variant_of link for a tracked variant), *status* over a declared lifecycle (at minimum: **probationary**, active, rolled back, suspended — a new version enters probationary and *earns* active under §11's evidence rule, and rolled-back-by-Steward and suspended-by-Auditor stay distinct, attributable acts), and *activation time*; and every recorded act carries **who acted** — role, version, *mode* (agent, human, or program), and the **principal**: the authenticated identity behind the act — the named human for a human-held act *in any capacity* (HB-3, §14), the agent or program instance identity otherwise — plus the Office when a human acts in Board capacity. Impersonation-binding (§3), separation of record (§16), and named-human attribution (HB-3, §17) are enforceable only if identity *and* capacity are recorded per act. A per-version scorecard (runs, pass rate, attributed cost) is then derivable from the event history rather than separately maintained. This is what lets a flow pause on Friday and resume on Monday, and what lets a human take over a role with full context. For a takeover to be more than picking up raw state, this history must capture the *decision trail* — what was decided and why, not only what changed — so whoever inherits the role inherits the reasoning, not just the outcome. Legible reasoning, not merely stored state, is what makes a clean handover possible. Two guarantees keep the trail trustworthy. **Provenance:** stored reasoning is *trusted* by whoever inherits it, so entries and summary spans derived from untrusted external content stay labeled as such, and a resumed flow — or a human reading a briefing — applies the same data-not-instructions posture to replayed history as to fresh tool output; without this, the memory plane launders an injection into trusted state (§14). **Decisions are consumed, never re-made:** any decision that shaped a run — the Optimizer's routing choice included, not only human injections — is recorded in the event history and reused on resume, or the replay guarantee of §6 silently breaks.\n- **Observability plane** — full-trace capture of every step, tool call, decision, cost, and output for every role. Not event logging — session-level trajectories. The raw signal the Steward, Optimizer, Auditor, and Verification consume. Capture is **mediated, not volunteered**: traces are recorded by the plane at the runtime and tool-call boundary — the substrate the role runs on, not the role's own code — and a role's self-emitted signals are claims to be cross-checked against mediated capture. Any signal consumed for an authority-relevant decision (a suspension, a rollback, an autonomy raise) must come from a source the measured role cannot write; an assurance stack fed by the self-reports of the parties it measures is not an assurance stack.\n- **Control plane (the Handbrake)** — the human-with-AI-skill interface for pausing, inspecting, adjusting, and resuming any flow. Detailed below; it is the architectural centerpiece.\n\n**Data governance across the planes.** Full-trace capture will contain client data, secrets, and personal data, so the planes carry it as governed content — stated, as ever, as guarantees: events and artifacts carry a **classification**; a role contract's *Inputs* bound what that role — and any human impersonating it — may read from the planes (least privilege applies to plane reads, not only to tool access); the constitution declares **retention and erasure** policy, and erasability must coexist with tamper-evidence (payloads erasable, integrity chain preserved — mechanism open); and data leaving the cell to an external implementer or model provider is an outbound boundary crossing, governed like any other external effect (§8).\n\n---\n\n## 6. The Handbrake (control plane in depth)\n\nThe handbrake is the debugging mode of the enterprise. The analogy is exact: a car has a maintenance\u002Fdebug mode where a technician halts normal operation, inspects and tunes internals, then returns the car to normal mode. The same must be true of every flow in the organization.\n\nIt is built from five primitives that already exist as engineering patterns — checkpointing, breakpoints, and record-replay are proven as-is (this is the durable-execution pattern, §19); mid-flow injection into agent state is the same patterns applied to a newer substrate:\n\n1. **Breakpoints** — declared pause points, set *before* or *after* any step. Some are static (always pause here — e.g. before an irreversible action); some are dynamic (pause only when a condition or confidence threshold is met).\n2. **State inspection** — at a breakpoint, the full state is readable: what has been done, what was decided and why, what it cost, and exactly where it paused. A human taking over should receive this as a readable briefing — a summary of recent activity and the exact decision point — not raw state alone; the handbrake is only as useful as the human's ability to understand what they are looking at.\n3. **Injection** — the human does not just approve or reject. They can supply an edited output, missing context, a corrected decision, or a direct CLI\u002Fprompt-level instruction that overrides what the agent was about to do.\n4. **Resume** — the flow continues from the exact paused point, consuming the injected value instead of re-deciding. It does not restart from the top.\n5. **Replay** — any past run can be reconstructed step by step *from the recorded trail* — inputs, outputs, and the decision trail (§5) — to find where a bad decision entered (\"time-travel debugging\" in the record-replay sense). Replay is reconstruction, not re-execution: an LLM step is not reproducible in general (inference is nondeterministic; provider-side updates silently change the function being called), which is exactly why the version registry pins the behavioral bundle and the trail must capture *why*, not only *what*.\n\n**Hard requirements that make the handbrake real (HB-1…HB-4):**\n\n- **HB-1.** A **durable checkpointer** must persist exact state at every meaningful step, or pause\u002Fresume is impossible.\n- **HB-2.** Every tool\u002Faction must be **as safe to retry as its effect allows** (INV-4): idempotent where the effect is owned or reversible; an at-most-once *attempt* with a recorded outcome — plus compensation where reversal exists — where it is not. Resume never re-fires an effect whose prior attempt is recorded, and never skips one that did not happen.\n- **HB-3.** The handbrake must be present on **every** flow as a structural property, callable by any authorized human at any time — not a special path some flows happen to support. And *authorized* is a governed word, not a vibe: handbrake access presupposes authenticated, attributable human identity (§14, trust boundaries); the authorized-humans list is Governance-plane data, scoped per role and autonomy level; changes to that list — and to the break-glass roster (§17) — are high-blast-radius acts under §8; and every injection is attributed to a named human in its own audit trail.\n- **HB-4.** **Re-entry is resume, not restart.** Re-entering a flow after implementer death behaves like resume: a completed flow idempotently returns its recorded outcome with no new events; a crashed one continues from its last durable step; recorded decisions are consumed, never re-made (§5).\n\nThis is what makes \"a human jumps in, adjusts, then steps back out\" a first-class operation rather than an emergency.\n\n---\n\n## 7. Agent-first, human-tolerant execution\n\nProcesses run at agent latency by default. The hard problem is graceful degradation when a human assumes a role: a human is a slow node, and the system must absorb that without stalling.\n\nMechanisms:\n\n- **Asynchronous, event-driven coordination.** Roles communicate through durable messages and events, not blocking calls. A waiting human holds up only the work that genuinely depends on their output.\n- **Durable pause\u002Fresume.** A flow waiting on a human is a paused, checkpointed flow — not a thread burning resources. It can wait minutes or days at no cost.\n- **Buffering and backpressure.** Upstream roles keep producing into queues; downstream flow control prevents a slow human node from cascading stalls through the system.\n- **Elastic SLAs.** Service expectations flex by implementer mode. A goal's deadline accounts for whether a role is currently agent-held or human-held.\n- **Reroute where possible.** Independent work continues around the human node; only dependent work waits.\n\nThe principle holds cleanly for bounded, low-fan-out roles: a human there changes the *speed* of one node, not the *correctness* of the system. It is not universal. For a high-throughput agent role — an Orchestrator fanning out thousands of parallel tasks — a long human pause eventually saturates the buffers, and backpressure then reduces the *liveness* of the dependent subtree. So polymorphism guarantees a human can always *stop and inspect* any role; it does not guarantee a human can *run* one at its native throughput. The honest rule is takeover where the role is bounded, and suspension where it is not (INV-2; the mode is declared per contract, §2).\n\nThere is a second honest cost, managed rather than denied: moving humans out of the loop for throughput degrades their readiness to intervene — the ironies of automation (§19) — and the model's own success erodes the practice that keeps its safety valve competent. Takeover competence is therefore a *maintained* resource, not an assumption: the constitution declares per-role takeover drills on live or replayed flows (the Replay primitive is a ready-made simulator), and time-to-competent-intervention is an observable the Steward reports and the Board reviews (§3, §12).\n\n### Why the gates don't slow it down\nA reasonable worry is that a Board, a Governance plane, an Auditor, and a Verifier add up to gridlock. They do not, because the human-speed functions sit *outside the hot path*: the Board is constitutional and periodic, never in the loop; the Auditor normally only monitors, asynchronously; Governance is compiled runtime checks, not meetings; and the only *always*-inline gate, Verification, runs at agent speed (L0\u002FL1 breakpoints are also inline, by design — but only for high-blast-radius action classes). The model multiplies *governance concepts*, not *synchronous human approvals* — which is exactly what lets it stay faster than a management hierarchy while being better governed. To keep it that way, gates are reviewed periodically and any that no longer earn their place are retired — governance is pruned, not only added, so checks cannot quietly accumulate back into the hot path. Gate *health* is monitored the same way: an L1 gate whose approval rate saturates near 100% is either dead weight or a rubber stamp — approval fatigue is a measured effect, not a hypothetical — and either way it is redesigned, not left to decay.\n\n---\n\n## 8. Authority and autonomy model\n\nEvery action class carries a declared autonomy level. Blast radius — the size and reversibility of consequences — sets the ceiling.\n\nBlast radius is two axes, not a feeling: **reach** (who is affected if this goes wrong) and **reversibility** (what undoing costs). The class takes the *worse* of the two — a trivially wide action and an irreversible narrow one are both high-blast — and when the two are in doubt, round up; that is the same fail-safe posture the novel-action rule below applies to the unclassified.\n\n| | Undo is cheap (minutes, yours) | Undo costs real effort | No undo exists |\n|---|---|---|---|\n| **Reach: inside the cell** | L3 | L2 | L1 |\n| **Reach: other cells \u002F the org** | L2 | L1 | L1 |\n| **Reach: outside world (clients, public, regulators)** | L1 | L1 | L0 |\n\nThe table is a *default mapping*, not a verdict — a cell's constitution may only tighten it (assign a lower level), never loosen it. It exists so \"high blast radius\" is an answer to two checkable questions rather than a judgment call made under deadline.\n\n| Level | Behavior | Use for |\n|---|---|---|\n| **L0 — Suggest** | Proposes; a human acts. | Highest-risk, irreversible actions. |\n| **L1 — Act with approval** | Prepares the action; pauses at a breakpoint for human approval before it takes effect. | High blast radius, reversible with effort. |\n| **L2 — Act and report** | Acts within policy, then reports for after-the-fact review. | Routine, low blast radius. |\n| **L3 — Fully autonomous** | Acts within policy, no per-action review; subject to monitoring. | Well-understood, low-stakes, high-volume. |\n\nRules: autonomy is assigned **per action class, not per role**; the same role may operate at L3 for safe actions and L0 for dangerous ones. Autonomy is raised over time as trust is earned from observed performance, never granted by default — but because raising an action class *is* a change to enforced governance, an agent never raises its own. The Observability plane and the Auditor *propose* an increase as a machine-surfaced amendment (the learning loop, §17); a human ratifies it. Performance earns a proposal, not an automatic promotion — the lever is always pulled by a human (invariant #10). Higher autonomy always implies stronger monitoring, not weaker. This model also sets the **capability floor** the Optimizer (§10) must respect. Risk classes are coarse and inherited by default — an action takes the class of the category it belongs to, refined only where blast radius actually varies (invariant #8) — so classification is a small governed set, not a per-action burden across thousands of cases. The hard case is the genuinely novel action — one with no prior class, like the new regulated capability in the worked example (§18). The model's answer there is fail-safe, not fast: an unclassified novel action inherits the *highest* risk class by default (lowest autonomy, human-gated) and at the same time raises a classification proposal to the Board. This is deliberate about a real cost — genuinely novel high-blast-radius work *is* slower the first time, because treating the never-before-done as low-risk to keep up speed is precisely the mistake the model exists to prevent. Once classified, the class is reusable and fast thereafter.\n\nThree refinements close loopholes this model would otherwise leave open:\n\n- **Field evidence has a statistical floor.** Observing zero failures in *n* runs bounds the failure rate only below roughly 3\u002F*n*, so telemetry alone can earn raises for high-volume, low-severity classes — while a class whose feared failure is rare and severe requires complementary evidence (targeted adversarial evaluation, staged canary exposure under a capped blast radius) before a raise is ratified.\n- **Deploying a new version into a high-authority role** (Direction, Orchestration) is itself a classified action with a blast-radius level — who may deploy is governed, not implicit — and the new version enters the registry as *probationary* (§5, §11).\n- **The humans at L0\u002FL1 gates act through the constitution channel.** Approval and execution are *gate-powers* the constitution explicitly grants to humans (INV-9, §17) — Office-acts on the human side of the boundary, clause-traced and separately audited. A human *executing* an L0 action is not impersonating the Role, whose authority is suggest-only; they exercise a granted power — which is how the impersonation-binding rule and L0 coexist.\n\n---\n\n## 9. The Steward (the \"org doctor\", named for what it does)\n\nA deliberate check-and-balance: the most autonomous, highest-authority roles (Direction, Orchestration) are the most dangerous if they drift, hallucinate, or degrade. So the model places a **maintenance function over them that can fix them but cannot use their power.**\n\n**Name.** *Steward* (full: **Reliability & Conformance Steward**). Alternatives if you prefer a different flavor: *Homeostat* (keeps the system in equilibrium), *Diagnostician*, *Conformance Warden*, *Org Reliability Engineer*. The chosen name should signal maintenance and health, not command.\n\n**Responsibilities — on two distinct objects, named per act:**\n- Watch the behavioral health of the other roles: drift from goals, hallucination, degraded quality, looping, runaway cost, policy violations.\n- Contain the running **flow**: quarantine a flow that loops, spirals on cost, or approaches its budget cap, and roll it back to a known-good checkpoint — the continuous enforcement arm of §14's budget and loop guardrails.\n- Maintain the live **instance**: adjust configuration and prompts, reset or restart from a known-good checkpoint, and fail over *provisionally* to a registered known-good implementer or version already authorized for the Role (tracked in the registry, auto-flagged for review). *Adopting* a different implementer permanently is the Board\u002Fhuman path of §17, which the Steward cannot take.\n- Act on the Auditor's regression alerts (§11): when a version is rated a regression, the Steward is the role that actually rolls it back.\n- Quarantine a drifting role-holder and flag for human takeover before it causes harm.\n\nThe two objects have different blast radii and different rollback targets — a checkpoint in a flow's history versus a version in the registry — and the record names which breaker fired. The Steward's own writes are graduated like everything else (§8): emergency *containment* — quarantine, rollback to known-good — is unilateral, because pause is safe; *rewrites* of a high-authority role — retuning Direction's prompts or configuration, a provisional failover of Direction or Orchestration — are prepared-and-approved acts (L1, or a two-key rule), because rewriting the Director is the very mechanism by which business behavior is shaped, and \"no business decisions\" must be enforceable, not aspirational. §14 names Steward compromise accordingly.\n\n**Hard boundary — the Steward may not:**\n- Make or change any business decision, priority, or strategy.\n- Approve work products or act in place of Verification.\n- Override the Governance plane.\n\nIt is the difference between the engineer who can restart, tune, and roll back the leadership system, and the executive who decides what the company does. One keeps the machine running correctly; the other decides where the machine goes. The Steward is strictly the former. It governs *operational* drift of *live instances*; *purpose* drift is the Board's job (§3) and *version* fitness is the Auditor's (§11).\n\n**Mode.** Agent by default (continuous monitoring is an agent-scale task), with human impersonation on demand for hard maintenance calls and audits — the same polymorphism as every other role.\n\n---\n\n## 10. The Optimizer (capability-to-task matching)\n\nA non-authoritative system role, sibling to the Steward: the Steward optimizes for **reliability**, the Optimizer for **efficiency and capability-fit**. Neither decides what the org does.\n\n**What it does.** Sits optionally between steps and matches each task to the **minimum-capability implementer that still clears the task's risk and quality floor**. This is not cost-minimization — it is capability-to-task matching, and it sometimes spends *more*. Examples, software and non-software alike:\n- Trivial, low-stakes task (update some links; draft a routine acknowledgment) → a light, cheap model.\n- Moderate task (add an integration; summarize a contract) → a mid-tier model.\n- Novel, high-stakes task (design a new authentication method with access rights; draft and negotiate a binding clause — a human Office executes it, §17) → the strongest model.\n\n**The hard constraint.** The Optimizer is **bounded by the autonomy\u002Fgovernance model (§8)**: the task's risk class sets a capability floor, and the Optimizer may only minimize cost *beneath* that floor. A high-blast-radius task may never be routed to a weak implementer to save money — that is not frugal, it is dangerous. Safety and quality are never traded for cost. Critically, the risk class is **constitutional input, not an Optimizer judgment**: the Optimizer optimizes beneath a floor it does not set. A role that could both classify a task's risk *and* optimize against it could quietly lower the floor to save cost — so classification is governed (§8, §17) and the Optimizer's routing is itself auditable.\n\n**The loop to guard.** There is a subtler way the floor can erode: floors are ratified by the Board, and the Board reads telemetry — much of it produced by the Optimizer itself. If Optimizer output can shape the floor it optimizes against, the role has indirectly authored its own constraint, which is invariant #10 failing in slow motion. The model closes the loop the same way §8 handles autonomy raises: telemetry may *inform* a floor proposal, but a floor change is an amendment — machine-surfaced with its **provenance visible** (\"this proposal originates in Optimizer cost data\"), ratified by humans who can see the source's incentive, and re-validated at compilation. Data earns a proposal; only a human moves a floor — and never on the unexamined word of the role that profits from the move.\n\n**Where it gets \"best version\".** When it must pick among versions of an implementer, it consumes the **Auditor's version ratings (§11)** — it does not judge version fitness itself; it routes to the version the Auditor currently rates fittest for the task class. Eligibility, though, is registry **status**, not rating: the Optimizer ranks only among status-active versions, and status changes (a suspension, a rollback) are events consumed before any subsequent dispatch — because ratings are accumulated signals and necessarily lag the breaker (§11).\n\n**Three edges, defined.** When *no* candidate clears the floor, the Optimizer escalates (§12) — it never relaxes a floor to proceed. Before attributed history exists, routing runs on declared nominal capability and cost until the record accrues; a *probationary* version is routed only within its probationary bounds (§5, §11). And the floor may carry a **diversity constraint**: for high-blast-radius classes the constitution may require the Verifier seat to be implementation-independent of Execution (§4.4) — a constraint the Optimizer enforces in routing like any other floor.\n\n**Where to place it (YAGNI).** The Optimizer itself consumes latency and tokens. Insert it only between steps where the cost-or-capability spread is wide enough to pay for the routing decision. A uniform pipeline does not need one — routing it is pure overhead.\n\n**Mode.** Agent by default, human impersonation on demand.\n\n---\n\n## 11. The Auditor (version fitness and safety)\n\nAgents ship in versions, and several versions or parallel variants may run at once — Executor v2 alongside v3, Optimizer v5 under trial. Someone must judge, on the fly and across accumulated activity, whether a version is getting better, getting worse, or becoming dangerous. That is the Auditor. Its human analogue is a **periodic audit team — but run as agents, for agents, continuously.** It does not operate, steward, or direct. It audits, rates, and reports.\n\n**What it judges (the object nothing else owns).** The **version**, evaluated as a population over time. This is distinct from the Verifier, which scores a single *output*, and the Steward, which watches a single *live instance* in real time and can fix it. The Auditor's object is the release itself, judged across many runs.\n\n**Responsibilities:**\n- Track every agent version and parallel variant via the version registry (§5).\n- Rate versions from accumulated field activity — quality, cost, latency, drift, incident rate — producing a fitness rating \u002F leaderboard per role.\n- Detect regressions (a new version performing worse than its predecessor) and dangerous behavior.\n- Normally: **monitor and report only.**\n- On a dangerous or severely drifting version: **suspend it** (a circuit breaker) and **escalate to humans**.\n\n**Hard boundary — the Auditor may not:**\n- Modify, retune, or rewrite an agent — that is the Steward.\n- Dismiss or retire a version — that is a human \u002F Board decision.\n- **Reinstate** what it suspended — lifting a danger-grade suspension requires a *human decision*, recorded in the audit trail, resolving the escalation; the Steward may *execute* the reinstatement, never decide it. (Ordinary regression alerts are different: there, Steward-autonomous rollback and restore is the designed path.) *Pause is unilateral (safety); un-pause is a human act.* No agent — and no pair of agents — suspends and then quietly un-suspends.\n- Make any business decision.\n\n(No burning at the stake: the Auditor can stop a version to prevent harm, but it cannot condemn, alter, or revive one.)\n\n**Why it is not redundant — it closes the evaluation loop.** Two existing system roles silently assume a signal that no role produced: the Optimizer assumes it knows which version is \"best,\" and the Steward assumes it is told when a release regressed. The Auditor is that signal:\n\n```mermaid\nflowchart TD\n    V[\"Per-output scores\u003Cbr\u002F>(Verifier)\"]\n    OBS[\"Traces: cost \u002F latency \u002F drift\u003Cbr\u002F>(Observability plane)\"]\n    EVT[\"Incidents\u003Cbr\u002F>(events)\"]\n    AUD[\"AUDITOR\u003Cbr\u002F>Rates agent VERSIONS over accumulated activity\u003Cbr\u002F>→ fitness leaderboard · regression + danger detection\"]\n    OPT[\"OPTIMIZER\u003Cbr\u002F>Routes to the fittest version\"]\n    STW[\"STEWARD\u003Cbr\u002F>Rolls back \u002F retunes the live instance\"]\n    HUM[\"HUMANS \u002F BOARD\u003Cbr\u002F>Reviews suspended version\u003Cbr\u002F>A human decides reinstatement —\u003Cbr\u002F>the Steward may only execute it\"]\n    WARN[\"The Auditor may unilaterally SUSPEND\u003Cbr\u002F>a dangerous version — never reinstate it.\"]\n\n    V --> AUD\n    OBS --> AUD\n    EVT --> AUD\n    AUD -->|ratings| OPT\n    AUD -->|regression alert| STW\n    AUD -->|danger| HUM\n    HUM --- WARN\n\n    style AUD fill:#1a3a6a,color:#fff,stroke:#0d2244\n    style WARN fill:#fff3cd,stroke:#856404,color:#533f03\n```\n\n**Breaker precedence (deconfliction with the Steward).** Two circuit breakers on two different objects: the **Steward** pauses-to-fix a *live instance* (or a running flow, §9) and may restart it as part of the fix; the **Auditor** suspends-and-escalates a *version* it has no authority to touch. But every live instance runs *some* version, so the axes join — and the join is governed, not hoped away. **What suspension does, mechanically:** it compiles into a Governance-plane predicate, enforced at the action site like any other rule (§17) — no new dispatch to the suspended version, and in-flight actions of that version are blocked at the pre-effect check or checkpointed at the next step. The breaker is enforced by the plane, not by the Steward's reaction time. **Precedence is explicit:** an Auditor suspension outranks a Steward restart for the affected version — the Steward may migrate work to another version; it may not re-activate the suspended one (see the reinstatement rule above). Independent breakers on independent axes, with a defined rule at the one point the axes meet.\n\n**The suspension threshold is governed, not improvised.** Suspension is a heavy act, so the bar is set in the constitution, not left to the Auditor's discretion: it is reserved for *danger* — behavior that risks harm — while ordinary regressions are alert-only, routed to the Steward and Optimizer rather than suspended. To prevent a circuit-breaker cascade, suspensions are rate-limited and one suspension may not auto-trigger others; and because the Auditor cannot reinstate, every suspension carries a **defined human-response SLA** declared in the constitution, so a suspended-but-critical version cannot hang indefinitely waiting on no one. A missed SLA is itself a governed event, not a silent hole: it escalates up the Office ladder and, if still unanswered, onto the break-glass path (§17), so a stuck suspension always surfaces to an accountable human. And if *no* human answers at all, the terminal behavior is defined rather than circular: the suspension holds and the cell degrades to its constitution-declared **safe mode** — new work in the affected class pauses; nothing waits, unbounded, on a human who is not coming.\n\n**Evidence floors, honestly stated.** Below a constitution-declared minimum of accumulated activity a version is rated **unproven** — not judged for fitness — with one deliberate exception: *danger detection is exempt from the evidence threshold*; a safety breach is not an evidence-quantity question. Symmetrically, the Auditor's power is stated honestly: it detects cost, latency, and pass-rate regressions strongly, and rare dangerous tail behavior weakly — the suspension breaker is defense-in-depth, not a reliable dangerous-version detector, which is why §8 demands complementary evidence (adversarial evaluation, capped canaries) exactly where field telemetry is structurally insufficient.\n\n**When to switch it on (invariant #8 \u002F YAGNI).** Only once you actually run multiple versions or parallel variants. With one version per role there is nothing to compare yet — turn it on when versions start flying.\n\n**Mode.** Agent by default, human impersonation on demand.\n\n---\n\n## 12. Escalation and human takeover\n\nA role escalates — pauses and requests a human implementer — when any of these fire:\n\n- Confidence falls below the threshold for its current autonomy level.\n- The situation is out-of-distribution: novel, ambiguous, or unspecified.\n- The action would exceed the role's authority scope.\n- Governance flags a policy boundary.\n- The Steward detects drift and quarantines the role, or the Auditor suspends a dangerous version.\n\nEscalation is a clean substitution: the flow checkpoints, a human assumes the role's interface (impersonation on demand), acts or corrects, and either hands the role back to an agent or stays for the duration. Because state lives in the memory plane and the contract is fixed, the takeover requires no special wiring. A human entering a role this way is bound by the same governance as the agent (INV-9, §3). Four rules make the substitution exact:\n\n- **Every escalation runs against a declared response.** The constitution states, per role, the escalation roster and its response SLA — §11's suspension SLA is one instance of this general rule, not an exception — and a missed SLA escalates the same way: up the Office ladder, onto the break-glass path, then the declared safe mode. A paused L1 flow never waits on nobody, indefinitely, by omission. The roster itself is *resourced*: standby duty is rostered, drilled (§7), and compensated as the constitution provides (§17) — a hat with no roster is how takeover fails at 2 a.m.\n- **A takeover opens a tracked variant** in the registry for its duration (§5). When the trigger was an Auditor suspension, the human runs as a fresh variant whose lineage names the suspended parent — the suspended version itself never acts.\n- **The Auditor's suspend power reaches human-held variants too.** INV-9 means a seat is bound by the Role's constraints regardless of who fills it — with immediate notification to the impersonator's Office; ejecting a human from a seat is exactly as loggable, and exactly as escalated, as suspending an agent.\n- **For a role whose contract declares suspend-and-inspect** as its human-takeover mode (§2, INV-2), escalation means: the flow checkpoints, the role suspends, a human inspects and corrects through the handbrake, and an agent implementer resumes — the human corrects the role without pretending to run it.\n\n---\n\n## 13. Reference topology\n\n```mermaid\nflowchart TD\n    BOARD[\"REPRESENTATION & ACCOUNTABILITY — The Human Board\u003Cbr\u002F>Offices: Board \u002F Chair \u002F CEO \u002F CFO … (one person or many)\u003Cbr\u002F>Writes & maintains the CONSTITUTION · Periodic alignment review\u003Cbr\u002F>Legal & public accountability · May mandate change\"]\n\n    subgraph PLANES[\"THE FOUR PLANES — cross-cutting: every role below runs INSIDE all four\"]\n        direction LR\n        GOV[\"GOVERNANCE\u003Cbr\u002F>Compiled policy + audit trail\u003Cbr\u002F>checked on every action\"]\n        MEM[\"MEMORY \u002F CONTEXT\u003Cbr\u002F>Event history · decision trail\u003Cbr\u002F>version registry\"]\n        OBS[\"OBSERVABILITY\u003Cbr\u002F>Mediated full-trace capture\u003Cbr\u002F>cost attribution\"]\n        CTRL[\"CONTROL — the Handbrake\u003Cbr\u002F>Pause · Inspect · Inject · Resume · Replay\u003Cbr\u002F>on ANY role, at any time\"]\n    end\n\n    subgraph CHAIN[\"THE VALUE CHAIN — agent speed\"]\n        DIR[\"DIRECTION\u003Cbr\u002F>What & why · Director\"]\n        ORCH[\"ORCHESTRATION\u003Cbr\u002F>Who & when · Orchestrator\"]\n        EXEC[\"EXECUTION\u003Cbr\u002F>Executor\"]\n        VERIF[\"VERIFICATION\u003Cbr\u002F>Verifier\"]\n        OPT([\"Optimizer — optional, inserted between steps\u003Cbr\u002F>where the capability\u002Fcost spread pays for it\"])\n    end\n\n    subgraph SYS[\"SYSTEM ROLES — no business authority\"]\n        STWD[\"STEWARD\u003Cbr\u002F>Health of live roles & flows\"]\n        OPTIM[\"OPTIMIZER\u003Cbr\u002F>Capability-to-task matching\"]\n        AUDIT[\"AUDITOR\u003Cbr\u002F>Rates versions · may suspend + escalate\"]\n    end\n\n    BOARD -->|\"constitution — compiled into the Governance plane (§17)\"| PLANES\n    PLANES -->|\"span and govern every role\"| CHAIN\n    PLANES ---|\"span and govern\"| SYS\n    DIR -->|\"prioritized goals\"| ORCH\n    ORCH -->|\"decomposed work\"| EXEC\n    EXEC \u003C-->|\"produce → score → revise\"| VERIF\n    OPT -.->|\"routes implementer\u002Fversion\"| EXEC\n\n    style BOARD fill:#1a3a6a,color:#fff,stroke:#0d2244\n    style PLANES fill:#e8f0fe,stroke:#2d5086,color:#1a3a6a\n    style GOV fill:#2d5086,color:#fff,stroke:#1a3a6a\n    style MEM fill:#2d5086,color:#fff,stroke:#1a3a6a\n    style OBS fill:#2d5086,color:#fff,stroke:#1a3a6a\n    style CTRL fill:#3a6aa0,color:#fff,stroke:#2d5086\n    style CHAIN fill:#f8fafc,stroke:#2d5086,color:#1a3a6a\n    style DIR fill:#2d5086,color:#fff\n    style ORCH fill:#2d5086,color:#fff\n    style EXEC fill:#2d5086,color:#fff\n    style VERIF fill:#2d5086,color:#fff\n    style OPT fill:#d4edda,stroke:#28a745,color:#155724\n    style SYS fill:#f0f4ff,stroke:#2d5086,color:#1a3a6a\n```\n\nHumans hold Offices · agents fill Roles · any Role can be impersonated on demand — and every role, operating and system alike, runs inside all four planes (§5); the planes are not stages in the chain.\n\n---\n\n## 14. Failure modes and the guardrails that contain them\n\nGrouped by type. As a rough rule, governance and organizational failures dominate during the transition into the model; operational and safety failures dominate once it runs at scale. The catalog is consistent with — and organized by role against — the empirically measured failure distribution of multi-agent systems (§19: MAST): specification failures land on Direction, inter-agent misalignment on Orchestration, verification failures on the Verifier gate.\n\n**Governance failures**\n\n| Failure mode | Guardrail |\n|---|---|\n| **Purpose drift** (org flawlessly does the wrong thing) | Board periodic human-interest alignment review (§3); constitution as fixed reference; binding mandate to correct. |\n| **Board becomes a shadow operator** (reintroduces the human bottleneck) | Board acts only through the constitution and bounded Role-impersonation, never turn-by-turn; changes are audited amendments. |\n| **Constitution is wrong, or in crisis** | Bounded, auto-expiring break-glass for emergencies plus a standing constitutional-review trigger; neither can change the constitution — only buy time until the Board amends it (§17). |\n| **Compilation drift** (enforced rules ≠ the written text) | Validation\u002Fattestation stage: every encoded rule traces to a constitution clause, re-validated on every amendment (§17). |\n\n**Operational failures**\n\n| Failure mode | Guardrail |\n|---|---|\n| **Role drift \u002F hallucination** in high-authority agents | Steward monitoring + quarantine; Verification gate; replay for root cause. |\n| **Version regression ships undetected** (new release quietly worse) | Auditor rates every version from field activity (§11); regression alerts to the Steward; ratings to the Optimizer. |\n| **Cost spiral** (agents loop, burn tokens\u002Fcompute) | Per-session cost attribution, budget caps, loop detection in the Observability","这是一个面向企业级AI应用的参考运营模型，定义了人与智能体协同工作的新型组织架构。核心特点是角色多态性（同一岗位可由人类、AI代理或程序实现）和全流程可中断性（所有自动化流程均支持人工暂停、干预与恢复）。模型采用分层设计：底层为执行角色的代理系统，上层为人类负责的战略决策与问责层（Board），中间通过细胞化单元实现灵活部署。适用于需要人机协同、合规可控、渐进式AI落地的各类组织，如金融服务、制造运营、科研管理等场景。",2,"2026-07-09 02:30:02","CREATED_QUERY"]