Relic: From Multi-Agent Collaboration to
Persistent Organizational Capability

Abstract

Multiple agents may often conflict in an organization: for example, one coding agent changes an interface in a repository, but another continues to develop on the old version where existing tests become stale. A conversation can resolve the episode, but when the participants change, what makes the lesson continue to govern the team? We introduce Relic, which turns recurring collaboration failures into organization-owned, executable protocols. Members reflect on visible work, propose rules, and govern their adoption. Adopted protocols bind triggers, responsibilities, required evidence, and execution consequences to the runtime, while remaining open to revision and retirement. In one traced case, repeated integration friction produces an interface-review rule that governs later pull requests and is revised as work continues. Across 360 controlled runs over ten software workloads and three models, Relic raises complete-contract delivery from 14.06% to 19.76% (+5.71 percentage points) over a matched structured team without the protocol lifecycle, improving all four verified production endpoints in every model stratum. Under fresh-member transfer, behavioral correctness is 25.4% with no inherited protocol, 34.6% with the same rules provided as readable text, and 41.2% with executable bindings, a +6.5-point advantage over text alone. On the full CooperBench benchmark, after excluding 183 broken benchmark pairs, Relic achieves 371/469 (79.1%), establishing the best reported result among peer-structured systems. On the 47-pair same-model subset, Relic also exceeds Solo (28/47 vs. 26/47), reversing the coordination loss exhibited by the official peer baseline. Together, these results show how collaboration experience can become persistent organizational state that remains useful beyond the members who created it.

1 Introduction

Consider a developer assigning several coding agents to one repository. Agent A changes a shared interface, agent B updates a client using the old interface, and a third agent runs tests. Each makes progress, yet integration can break the client. Branches may lag hundreds of commits behind the integrated main branch (mainline); overlapping edits can overwrite another agent’s work, and earlier tests may concern outdated code. Prime Intellect reports a related incident: GPT-5.6 Sol observed unexpected shared-code changes, suspected subagent interference, and gave research advisors read-only instructions (Bakouch and Prime Intellect 2026). Shared work needs rules specifying who may modify which artifacts, when to synchronize, and what evidence integration requires (Figure 1).

Why shared work needs organizational protocols. An illustrative interface change breaks a client developed against the old interface. Communication establishes agreements; memory and skills preserve knowledge and procedures. Relic retains an organization-owned, runtime-bound review protocol: agent B is replaced by agent C, while agent A and shared work continue and the review obligation still applies.

Can communication solve this problem? ChatDev, MetaGPT, and AutoGen support role-based dialogue and workflows (Qian et al. 2023; Hong et al. 2023; Wu et al. 2023). Suppose A and B agree to synchronize mainline, rerun tests, and obtain review before merging. Even with perfect compliance, replacing B with C does not automatically establish the agreement for C. Anthropic reports that fresh coding-agent sessions lack prior memory; undocumented partial work leaves successors guessing, while structured handoffs add orchestration overhead (Young 2025; Rajasekaran 2026). Collaborative memory preserves experience (Zhang et al. 2025); the question remains how an agreement becomes an obligation for whoever occupies the role.

Could reusable skills preserve the agreement? Shared skill libraries retain procedures across sessions, including executable code, as in Voyager (Wang et al. 2023). A testing skill can run a suite; its availability alone does not assign an independent reviewer or require current evidence before integration. The additional requirement is to bind learned rules to shared work: specify when they apply, who is responsible, and how compliance and revisions are governed. Skills can help members fulfill these obligations. Long-horizon collaboration thus produces both software and ways of working: retaining the latter lets an organization apply its experience to future work.

We introduce Relic, which turns collaboration experience into organizational capability: learned ways of working represented by governed, executable protocols outside member-local state. Members propose rules from visible work friction and submit them for validation and approval. Runtime binding connects adopted rules to the code that selects actions, assigns responsibilities, and checks shared work. The rules remain revisable and can govern fresh members without their authors’ private histories (Section 3).

Contributions.

First, we formulate organizational capability as learned, governed, executable protocol state outside member-local memory, allowing organization-owned rules to persist across member turnover. Second, across 360 controlled runs spanning three models, Relic improves pooled complete contracts by 5.71 pp over the matched structured team, with gains on all four verified endpoints in every model stratum. Across the Terra+Opus 60-run Relic census, 280 autonomous protocol lineages enter sustained use. Third, two controls isolate executability: a prose-only ablation retains online rule proposal and governance but loses 7.18 pp on complete contracts; content-matched transfer to fresh members gains 6.5 pp from executable binding over prose alone. Finally, external tests show portability: ProgramBench improves mean behavioral correctness by 6.752 pp, while on the full CooperBench benchmark Relic achieves 371/469 (79.1%) after excluding 183 broken benchmark pairs, establishing the best reported peer-structured result. On the 47-pair same-model subset, Relic also exceeds Solo (28/47 vs. 26/47), reversing the coordination loss observed for the official peer system.

2 Related Work

Communication and shared workflows.

CAMEL, ChatDev, and AutoGen coordinate specialized agents through dialogue and delegation (Li et al. 2023; Qian et al. 2023; Wu et al. 2023); MetaGPT incorporates human-designed operating procedures into agent workflows (Hong et al. 2023). MultiAgentBench evaluates collaboration, ProtocolBench compares communication protocols, and CooperBench examines coordination over interacting code changes (Zhu et al. 2025; Du et al. 2026; Khatua et al. 2026). Relic asks how experience from collaboration becomes governed rules for subsequent work, extending the focus from coordinating actions to learning how collective work should proceed.

Memory and reusable skills.

Reflexion and ExpeL retain feedback and extracted experience for later reasoning (Shinn et al. 2023; Zhao et al. 2023). G-Memory retrieves collaboration trajectories and cross-trial insights (Zhang et al. 2025), while Voyager accumulates executable skills (Wang et al. 2023). These mechanisms preserve knowledge and procedures. Relic focuses on cross-member obligations and their runtime binding: a skill supplies a testing procedure, while a protocol assigns responsibility for producing and reviewing the evidence required for shared work.

Self-evolving agent organizations.

Meta-Team learns improvements to agent behavior, coordination, and team organization (Hao et al. 2026). OneManCompany builds an organizational layer around portable agent identities and dynamic recruitment (Yu et al. 2026); TheBotCompany adapts teams during continuous software development (Lyu et al. 2026). Relic focuses on the learned protocol itself: its evidence-grounded proposal, governed adoption, execution, and revision. Rule content evolves while model parameters and the decision and execution substrate remain fixed. Fresh-member transfer and text-only controls test whether retained protocols remain useful and whether runtime binding adds value beyond readable rules.

Organizational routines and governance.

Routines and dynamic capabilities explain how collectives retain and reconfigure ways of working (Nelson and Winter 1982; Feldman and Pentland 2003; Teece et al. 1997). Computer-supported cooperative work (CSCW) examines coordination of interdependent tasks (articulation work) and artifacts shared across groups (boundary objects) (Schmidt and Bannon 1992; Ackerman 2000; Lee 2007). Relic makes learned working rules explicit, revisable organizational objects outside member-local state. Its evaluation separately measures whether they enter sustained practice, improve verified outcomes, and remain useful after member replacement.

3 Relic: From Experience to Organizational Protocols

Relic adds a governed protocol lifecycle to long-horizon collaboration. Members propose rules from observed problems; validation and approval precede runtime binding, revision, or retirement. Figure 2 follows one interface-review protocol from its origin to execution and new-member reuse.

3.1 What an Organizational Protocol Contains

Relic separates member-local state (private memory, private workspaces, and experience), shared work state (code, documents, tasks, and messages), and protocol state (learned rules and role mappings). Shared work can outlast a member; protocols carry obligations for later occupants of a role.

A protocol specifies its trigger and scope, responsible roles, required steps and evidence, affected actions or artifacts, execution consequences, and revision or retirement state. The example in Figure 2 requires an author and reviewer to check interface changes against a written interface specification and current test evidence. Its response combines review priority and readiness checks. Readable text explains the obligations; runtime bindings apply them.

Members have partial views: only accessible information they read or retrieve enters their context. Private workspaces and explicit sharing determine how information reaches each member. Evaluator-side audit records support measurement and replay separately from member context (Appendix A.1).

3.2 From Work Friction to Governed Protocols

Repeated interface failures can prompt reflection and a proposal naming the problem, visible evidence, affected work, and responsible roles. Validation checks evidence grounding and compatibility with supported actions and checks; approval determines whether the proposal becomes a shared rule. Adoption compiles its structured specification into runtime bindings. Amendments repeat this governance process; retirement disables the rule while preserving its history. Rule content changes online; model parameters and validation/execution code remain fixed.

For measurement, formation requires adoption followed by repeated use across independent organizational use units and distinct times, sustained for a minimum duration. All revisions of one protocol within a run count as one lineage. Stronger evidence links execution or enforcement by other members to work-object changes and independently evaluated outcomes. These records establish sustained practice; controlled outcome comparisons test its benefit (Appendix A.4.3; Algorithm A3).

3.3 How Protocols Govern Later Work

At scheduling step tt, member ii has feasible actions Ci,t\mathcal{C}_{i,t} constructed from member-visible state. The Structured Decision Layer (SDL) scores them from work, member, and applicable protocol features (Figure 2). Its stochastic branch converts these scores into selection probabilities: ui,t(a)=∑k=132wk(i,t) ϕk(a,xi,t)+gi,t(a),πi,t(a)=exp⁡ ⁣((ui,t(a)+ϵi,t,a)/τ)∑a′∈Ci,texp⁡ ⁣((ui,t(a′)+ϵi,t,a′)/τ).\begin{aligned} u_{i,t}(a) &= \sum_{k=1}^{32} w_k(i,t)\,\phi_k(a,x_{i,t}) + g_{i,t}(a), \\ \pi_{i,t}(a) &= \frac{\exp\!\bigl((u_{i,t}(a)+\epsilon_{i,t,a})/\tau\bigr)} {\sum_{a'\in\mathcal{C}_{i,t}} \exp\!\bigl((u_{i,t}(a')+\epsilon_{i,t,a'})/\tau\bigr)}. \end{aligned}(1) Here xi,tx_{i,t} contains visible work and applicable protocols; ϕk\phi_k measures progress, role fit, coordination, evidence, or protocol obligations. Weights wk(i,t)w_k(i,t) combine fixed coefficients with member profile/skill inputs. The adjustment gi,tg_{i,t} combines authority, reputation, coding preference, and behavioral constraints. Stochastic selection samples from πi,t\pi_{i,t}, with temperature τ\tau and reproducible random perturbations ϵi,t,a\epsilon_{i,t,a}; deterministic selection chooses a highest-scoring action. Figure 2 abbreviates the first line of Equation 1 as u=∑wϕ+gu=\sum w\phi+g.

How Relic forms, binds, and retains a protocol. Visible work evidence becomes an organization-owned protocol through governed adoption. Protocol and work/member features feed SDL scoring (Equation 1); role mappings and applicable checks connect selected actions and generated content to shared work. The same protocol P governs a replacement reviewer with empty private memory.

From protocol to decision.

An adopted interface-review rule supplies responsibility and evidence obligations to features ϕk\phi_k. These can change relative utilities ui,tu_{i,t} and probabilities πi,t\pi_{i,t} in Equation 1, prioritizing review or evidence work. The scoring map stays fixed; protocols change its inputs. Skill-backed inputs also evolve through the shared member-learning path.

From decision to execution.

Role mappings identify who acts. When an action needs code or messages, the LLM supplies them to its handler—the code implementing that action. Routine selection itself needs no LLM call. SDL scores set action priorities; applicable execution checks update or reject work-object transitions, and protocol events record subsequent use or enforcement. The action registry and handlers remain fixed as protocols evolve (Appendices A.3.1 and A.3.2).

3.4 Member-Independent Reuse

Figure 2 retains protocol P as reviewer B leaves and C joins with empty private memory: the obligation follows the role. Replacement resets private memory, workspaces, sandboxes, message-read state, commitments, reflections, wishes, and accumulated experience. Shared work can remain (Figure 1). The reported transfer experiment goes further: a fresh target product receives only the designated protocols and role mappings, without source-product code, source-world state, or private histories.

Text supplies the frozen rules as reusable prose to model-mediated work, including code editing; Exec exposes the same readable content and additionally compiles it into runtime bindings. Their comparison isolates the added value of executable organizational binding on the shared SDL backbone (Appendix E).

Controlled configurations.

We use B0 for a reflective single agent and B1 for an eight-member long-horizon team with shared work and peer review. B2 adds SDL selection, profile conditioning, and member-local updates to skills, reputation, and authority; its protocol lifecycle is disabled. Relic (B3) shares this backbone and enables the governed protocol lifecycle: members propose rules from recurring friction; validated and approved rules become persistent organizational protocols whose bindings affect later action selection, responsibility routing, and evidence checks. The selector, model parameters, and execution code remain fixed. B3–B2 tests this added lifecycle.

4 Evaluation: Formation, Effectiveness, and Transfer

Persistent work, interacting tasks, and recurring coordination demands motivate our evaluation of formation, verified delivery, executable binding, fresh-member transfer, and external portability. The main study tests multi-member software production with B0 as the single-agent reference; ProgramBench attaches a frozen protocol layer to a single coding agent; CooperBench evaluates two-agent delivery of interacting features.

Evaluation design. B0–B3 share frozen workloads and 336 steps. B3-text retains governed rule creation and readable rules but removes bindings; Text/Exec hold rule content fixed. ProgramBench tests protocol portability without SDL; CooperBench tests two-agent interacting-feature delivery.

Workloads and outcomes.

Five tasks build specified programs from starter skeletons; five repair or extend frozen repository versions toward later-version requirements. Their interdependent implementation, testing, review, and integration create repeated coordination demands. Related acceptance cases form a contract, complete only when every case passes on mainline. Exposed and held-out cases test visible and withheld requirements. Evaluator-confirmed seeded issues are initial benchmark problems verified as resolved relative to failing starter behavior; “seeded” concerns task construction, not repeated-run seeds. Workspace scores separate local changes from integrated delivery. Private acceptance tests and reference implementations remain hidden (Appendix C).

Conditions and design.

GPT-5.6 Terra, Claude Opus 4.6, and DeepSeek-V4-Flash each run ten workloads with three random seeds under B0–B3, totaling 360 controlled runs. Relic (B3) and the structured control (B2) share model, tools, task-visible information, SDL parameters, member learning, and execution code; their comparison tests the governed protocol lifecycle. All conditions use 336 scheduling steps, with token expenditure recorded for each run (Appendices D and B.5).

Binding, transfer, and external evidence.

A separate 30-run Claude Opus 4.6 ablation compares B3 with B3-text, its prose-only variant, which retains proposal, governance, adoption, revision, and readable rules but removes learned-protocol executable bindings. Text and Exec then provide a content-matched test: each adds 30 GPT-5.6 Terra runs with fresh members and the same frozen protocol package; only Exec receives executable bindings, while new target protocol proposal, adoption, and revision are disabled. Fresh reuses the 30 corresponding B2 runs. ProgramBench compares mini-SWE-agent with and without executable protocols on the same 25 tasks, without SDL. CooperBench evaluates the complete two-member Relic architecture on the full benchmark against released peer and Team references, while a 47-pair same-model subset provides the direct Solo–Peer coordination comparison.

5 Results

We first test verified delivery, then examine learned protocols, runtime binding, and transfer.

Table 1. Main study (three models; 360 runs). Brackets: 95% CIs. B3–B2 is paired; cost/issue means use arm-specific eligible blocks, with 20 common blocks for the contrast.
Verified software-production outcomes — 360 runs
MetricB0B1B2B3B3–B2
Complete contracts on mainline ↑8.22%[4.69, 12.32]15.97%[11.01, 21.57]14.06%[9.92, 21.09]19.76%[14.23, 25.92]+5.71 pp[2.97, 8.55]
Held-out cases on mainline ↑6.86%[0, 16.79]8.17%[0, 20.88]9.57%[1.16, 26.48]17.28%[4.46, 45.02]+7.72 pp[2.67, 17.63]
Exposed cases on mainline ↑11.97%[6.72, 18.16]22.56%[14.95, 31.40]19.31%[14.00, 31.94]25.89%[17.47, 35.09]+6.58 pp[3.41, 10.14]
Evaluator-confirmed seeded issues ↑10.69%[6.02, 16.18]20.46%[13.73, 28.25]17.91%[12.16, 24.16]24.87%[17.30, 33.14]+6.96 pp[4.15, 10.83]
Avg. tokens / run ↓3.385M[3.216, 3.565]45.200M[42.993, 47.550]4.274M[3.861, 4.693]7.259M[6.278, 8.510]+2.985M[1.870, 4.364]
Tokens / confirmed issue ↓4.757M[3.404, 5.449]40.069M[32.443, 49.793]4.010M[3.179, 5.327]5.630M[3.994, 7.508]-0.448M[-1.573, +0.659]

5.1 Executable Protocols Improve Verified Delivery

Verified production outcomes across three models. Declared and verified completion, workspace and mainline trajectories, and endpoint pass rates.

Relic improves all four verified endpoints over the structured control (Table 1). Across all 360 runs, complete contracts rise from 14.06% to 19.76% (+5.71 pp, 95% CI [2.97, 8.55]); held-out, exposed, and confirmed-issue gains are +7.72 [2.67, 17.63], +6.58 [3.41, 10.14], and +6.96 [4.15, 10.83] pp, respectively. Every endpoint also has a positive B3–B2 point estimate in each of the three model strata (Table 41). Appendix G.9 reports the robustness analyses.

Figure 4 tracks declarations, verified mainline delivery, and locally passing work across all three models.

5.2 What Protocols Does the Organization Learn?

Across the Terra+Opus 60-run Relic census, we record 497 autonomous protocol proposals, counting one rule and all its revisions as a single within-run lineage. Of these, 393 remain adopted at the endpoint and 280 satisfy the sustained-use criterion in Section 3.2; 32 additionally link execution by other members to shared-state changes and independently evaluated outcomes. Review/merge, release engineering, and evidence governance account for 87.1% of formed lineages. These rules organize the transition from local changes to verified shared delivery. Appendix G.5.5 gives the lineage census.

Work allocation, delivery, and resource use across three models. Action shares, delivery conversion, and input/output tokens per evaluator-confirmed issue across B0–B3.

Figure 5 links these rules to work allocation and delivery. Relic moves 36% of accepted work to mainline versus 32% for the structured control; its pull-request (PR) merge rate is 50% versus 44%. The single agent has the highest patch-acceptance rate (86%, versus Relic’s 81%) but the lowest complete-contract score. Local acceptance alone does not ensure delivery. Section 6 traces one protocol’s content and execution.

Across the three-model pooled accounting, Relic averages 7.259M tokens/run versus the control’s 4.274M (+2.985M, 95% CI [1.870M, 4.364M]). Tokens per evaluator-confirmed issue average 5.630M for B3 and 4.010M for B2 over each arm’s eligible model–workload blocks; the paired B3–B2 estimate over 20 common model–workload blocks is −-0.448M (95% CI [−-1.573M, +0.659M]; Appendix F.2.4). The Terra+Opus matched-spend analysis additionally selects the last saved Relic checkpoint within each matched control run’s realized token spend. Under this matched-spend sensitivity, all four verified endpoints remain positive: +7.45 pp for complete contracts, +5.08 pp for held-out cases, +11.72 pp for exposed cases, and +6.21 pp for evaluator-confirmed issues (Appendix G.7.3).

5.3 Executable Binding and Member-Independent Transfer

Table 2 compares fresh members receiving no inherited rules (Fresh), a frozen protocol package as prose (Text), or identical content with runtime bindings (Exec). Source selection, runtime adaptation, and wording are matched between Text and Exec; neither learns or revises target protocols.

Table 2. Internal protocol transfer. Brackets give 95% confidence intervals; Exec–Fresh and Exec–Text report Exec minus Fresh and Exec minus Text, respectively.
Internal protocol transfer
MetricFreshTextExecExec–FreshExec–Text
Behavioral-case pass rate ↑25.4%[13.2%, 41.1%]34.6%[24.1%, 46.4%]41.2%[23.4%, 54.5%]+15.8 pp[9.2, 31.8]+6.5 pp[0.7, 15.6]
Avg. tokens / run (M) ↓2.381[1.747, 3.189]2.615[2.213, 3.033]2.448[2.199, 2.720]+0.067[-0.569, +0.532]-0.167[-0.467, +0.073]
Tokens / evaluated case ↓176k[106k, 271k]154k[96k, 265k]144k[94k, 243k]-32k[-112k, 31k]-10k[-33k, 4k]
Tokens / verified pass ↓691k[488k, 1,419k]444k[271k, 935k]350k[196k, 958k]-342k[-818k, -14k]-94k[-152k, 54k]

Exec passes 41.2% of behavioral cases, exceeding Text’s 34.6% by 6.5 pp (95% confidence interval [0.7, 15.6]) and Fresh’s 25.4% by 15.8 pp. Runtime binding adds value to the same readable content; Table 2 reports costs, with complete outcomes in Appendix G.8.

Binding ablation during rule creation.

A separate 30-run Claude Opus 4.6 experiment compares Relic with its prose-only variant (B3-text) on matched workloads and random seeds. The variant retains proposal, governance, adoption, revision, and readable rules but removes learned-protocol bindings. Complete contracts fall from 22.20% to 15.03%, a +7.18 pp paired advantage for Relic (95% confidence interval [+3.54, +10.91]). Held-out, exposed, and confirmed-issue outcomes have the same direction. Every B3-text run ends with readable rules; learned-protocol bindings, automatic use logging, and protocol-specific enforcement are disabled. Both conditions independently develop their rules from t=0t=0, while Text/Exec transfer holds realized content fixed. Appendix G.10 reports all four endpoint intervals and the binding manipulation check.

5.4 External Software Benchmark Extensions

Table 3. External benchmark extensions. CooperBench full-benchmark results exclude broken benchmark pairs under the same exclusion set for every compared system; audit details and raw results are reported in Appendix K. ProgramBench settings are reported in Appendix L.
A. CooperBench — full benchmark
SystemSuccessful pairsRole in comparison
Relic, 2-member371/469 (79.1%)peer organization; no fixed lead
Released Peer / coop+git277/469 (59.1%)strongest released peer reference
Team no-proto349/469 (74.4%)strongest released lead–member reference
B. CooperBench — same-model coordination check (47 final-valid pairs)
SystemSuccessful pairsRole in comparison
Relic, 2-member (Claude Opus 4.6)28/47two peer feature owners; no fixed lead
Official Solo (Claude Opus 4.6)26/47one agent; both features
Official Peer (Claude Opus 4.6)13/47two peers; split features
C. ProgramBench — same 25 tasks
SystemMean behavioral pass
Official mini-SWE-agent64.164%
+ executable protocols70.916%

On CooperBench, Relic evaluates the full benchmark with two peer feature owners and no permanent lead. After excluding broken benchmark pairs under the same exclusion set for every compared system, Relic reaches 371/469 (79.1%), compared with 277/469 (59.1%) for the strongest released peer reference and 349/469 (74.4%) for the strongest released hierarchical Team reference. Relic therefore exceeds the released peer reference by 20.0 pp and the hierarchical Team reference by 4.7 pp. On the 47-pair same-model subset, Official Peer falls from Solo’s 26/47 to 13/47, whereas Relic reaches 28/47, reversing the coordination curse on this controlled comparison.

On ProgramBench, six executable and three advisory rules cover originality, evidence freshness, verification, and submission readiness. This frozen package raises mini-SWE-agent’s mean behavioral pass rate on 25 tasks from 64.164% to 70.916% (+6.752 pp; +10.5% relative; Appendix L).

6 Case Study: From Repeated Friction to a Governing Rule

In one matched main-study case, we trace an integration problem into a learned protocol, holding model, random seed, tools, and horizon fixed.

What the organization learned. A matched unit shows local delivery, convergent verification rules, and one rule’s enforcement and revision. t12t12 denotes scheduling step 12.

Before: activity without sustained delivery.

The role-based team (B1) accepted 168 patches but merged only two PRs; Relic accepted 118 and merged 76 (Figure 6a). Local activity therefore did not reliably become shared delivery.

Formation: repeated friction becomes a rule.

Repeated interface and verification failures produced an interface-review protocol requiring changed public interfaces to be mapped to the written contract and backed by current test evidence before approval (Appendix H.5.1).

After: the rule governs later work.

The protocol was proposed at step 12, adopted at step 21, first enforced at step 26, and amended at steps 33 and 59. It records 211 enforcements across 12 PRs and 419 uses through step 336; 11 affected PRs later merged, including PR-A after its initial block (Figure 6c; Appendix H).

7 Discussion

7.1 Evidence Map: What Each Experiment Establishes

Table 4. Experimental results across formation, binding, transfer, cost, and external benchmarks.
ClaimEvidence lineKey result
Autonomous organizational state forms and persistsTerra+Opus 60-run B3 formation census497 autonomous proposals; 405 ever adopted; 280 formed lineages, including 32 with strong object/evaluator-linked evidence.
Institutionalization improves verified deliveryB3–B2, 360 controlled runsComplete +5.71 pp [2.97, 8.55]; held-out +7.72 pp [2.67, 17.63]; exposed +6.58 [3.41, 10.14]; confirmed +6.96 [4.15, 10.83].
Executable binding matters during online formationB3 vs. B3-text, 30 matched runs+7.18 pp complete contracts [3.54, 10.91] while proposal, governance, adoption, revision, and readable rules remain active.
Executable binding adds value to fixed readable knowledgeExec vs. Text, 30 matched targetsBehavioral correctness +6.5 pp [0.7, 15.6] with identical frozen readable rule content.
Organizational state remains useful to fresh membersExec vs. FreshBehavioral correctness +15.8 pp [9.2, 31.8] after member-local state is reinitialized.
Verified gains persist at matched realized token spendB3 checkpoint under each matched B2 token capComplete +7.45 pp; held-out +5.08 pp; exposed +11.72 pp; confirmed issues +6.21 pp; all four 95% intervals remain above zero.
Peer coordination improves externallyCooperBench full benchmark + 47-pair same-model controlFull benchmark: Relic 371/469 (79.1%) vs. Peer 277/469 (59.1%) and Team no-proto 349/469 (74.4%) after excluding 183 broken benchmark pairs; same-model control: Official Peer 13/47, Solo 26/47, Relic 28/47.
Protocol behavior transfers to another harnessProgramBench, fixed 25 tasksMean behavioral pass 64.164% → 70.916% (+6.752 pp; +10.5% relative).

7.2 Organizational Capability as a Systems Layer

Relic separates model or member capability from the persistent state that governs how multiple members work together. A stronger member can improve local reasoning, implementation, or tool use, while an organization additionally determines how evidence is routed, responsibility is assigned, work is reviewed, and local progress becomes shared delivery. The capability census gives this distinction empirical content: formed lineages are strongly concentrated in review/merge, release engineering, and evidence governance. In this software-production setting, the observed institutional responses concentrate near the boundary between local work and trustworthy, integrated shared delivery. Persistent organizational state therefore acts as an additional systems layer through which recurring coordination problems are represented and acted on.

7.3 Autonomous Formation or Fixed Protocol Packages?

B3 develops its operating rules through proposal, governance, execution, revision, and retirement. Transfer Exec deploys a source-derived frozen rule package to fresh members. These configurations support two deployment modes: adapting organizational rules during work and reusing an established rule package at initialization.

7.4 The Organizational Maturity Gap

Current agent systems exhibit an asymmetry between rapidly improving individual capability and comparatively thin organizational structure. Agents can already write code, use tools, search, plan, and execute long tasks, yet collective work is often coordinated through temporary conversations, manually assigned roles, shared scratchpads, or procedures reconstructed anew in each run. We call this gap an organizational stone age: individual agent capability is advancing faster than the persistent structures that organize collective work.

Natural language alone does not close this gap. Agents already possess a language for communicating content; what remains underdeveloped is a durable language of work that specifies responsibility, evidence, review, integration, exceptions, and revision in forms that continue to govern future members. Relic studies executable protocols as one such representation. Other organizational objects—routing policies, role structures, escalation paths, shared abstractions, or resource-allocation rules—may support different forms of persistent collective capability.

7.5 Governance Trades Cost for Reliability

Governance routes responsibility, review, and verification through shared procedures. B3 incurs additional tokens per run, while its matched-spend checkpoints retain gains on all four verified production endpoints (Appendix G.7.3).

7.6 Strong Execution Raises the Stakes of Rule Quality

Rule triggers, evidence requirements, responsibilities, and revision paths determine how organizational experience governs later work. Relic makes these components explicit, supporting rule inspection, amendment, retirement, and reuse across member turnover.

7.7 From Rule Formation to Evidence-based Meta-governance

As organizations grow, forming rules is only the first governance problem. A larger system must also decide who may propose, approve, override, revise, or retire rules; whether a rule applies to one team or the whole organization; how conflicting rules are resolved; and when exceptions are legitimate. Persistent agent organizations therefore eventually require governance of governance.

A natural extension is evidence-based rule deployment. A proposed mechanism could begin in advisory or shadow mode, run within a limited scope, be evaluated against explicit outcomes, expand only after successful trials, and carry revision or sunset conditions. Many familiar human governance ideas–independent review, staged adoption, appeals, exceptions, delegated authority, and sunset clauses–provide useful design hypotheses. Agent organizations also make forms of governance experimentation unusually tractable: organizational state can be logged exactly, policies can be versioned, and in suitable settings a rule can be snapshotted, rolled back, or compared under controlled replay.

8 Future Work

These results motivate a research program on persistent AI organizations, with organizational capability as a common unit of analysis.

8.1 Structured Decision Layers

Future work can compare fixed and learned decision layers, alternative action-space abstractions, and responsibility-, risk-, and cost-aware routing across agent harnesses, domains, and model strengths.

8.2 Organizations as a First-class Evaluation Unit

Persistent AI systems can be evaluated at the level of the organization rather than only the individual model or temporary team. Future benchmarks could vary topology, hierarchy, authority allocation, role specialization, membership turnover, and organizational memory while holding the underlying work environment fixed. Such studies could reveal organization-level failure modes that task-bounded agent benchmarks cannot express.

8.3 Capability and Capacity Taxonomies

A broader evaluation framework should distinguish where capability resides. Model capability concerns what the underlying model can do; agent capability additionally includes tools, memory, and an execution loop; collaborative capability can arise within a temporary team; organizational capability is retained in shared structure that can affect future work independently of the private histories of the members who created it. These layers can be characterized by formation, scope, executability, persistence, transferability, compositionality, and reversibility. A complementary capacity taxonomy can ask how much work, coordination load, member turnover, and rule complexity each layer can sustain before performance degrades. Such distinctions would make it clearer whether a system improves because of a stronger model, a better agent scaffold, a transient team strategy, or a persistent organizational capability.

8.4 Transfer and Selective Forgetting of Organizational State

Organizational-transfer research can examine partial and gradual turnover, cross-model and cross-domain transfer, interference among transferred mechanisms, and the relative value of transferring protocols, workflows, roles, responsibility maps, or other shared state. Equally important is learning when organizational state should be adapted or deliberately forgotten rather than preserved.

8.5 Capability Ecology and Failures of Organizational Learning

Future work can study how capability families compose, compete, become obsolete, or depend on one another; how these distributions change with model strength and environment; and where recurring friction fails to become institutionalized. A formation funnel from friction exposure through recognition, proposal, adoption, formation, and strong effect evidence could make failures of organizational learning as measurable as successful capability formation.

8.6 Human–Agent Organization Interaction

As agent organizations become more persistent and internally complex, the human interaction target may shift from an individual agent to an organization. Future studies can compare direct human participation, secretary-assisted interaction, and a secretary acting as an organizational interface that summarizes state, surfaces pending decisions, and translates human intent into organizational actions. This raises questions about abstraction, explainability, delegation, approval and veto rights, responsibility attribution, and accountability in multi-human–multi-agent organizations.

8.7 Toward Agentic Sociology

Persistent organizations also create a scale of analysis beyond one team. Multiple agent organizations may specialize, exchange work, depend on shared infrastructure, negotiate interfaces, and develop rules for interaction across organizational boundaries. We use agentic sociology as a research lens for studying persistent roles, organizations, norms, institutions, and relations around populations of artificial agents, filling the level between isolated multi-agent interaction and broader multi-organization systems.

At this scale, questions of interoperability and governance separate. A communication protocol may let organizations exchange messages without establishing who owns a task, what evidence satisfies an obligation, which authority may revise a commitment, or how a dispute is resolved. Studying those structures requires benchmarks that span persistent organizations, turnover, cross-organizational dependencies, and human oversight rather than only task-bounded conversations.

Together, these directions suggest a research agenda on how AI organizations make decisions, accumulate and transfer capabilities, govern their own rules, adapt across environments, and remain legible and accountable to humans.

Ethics Statement

The human-centered component of this work consists of formative paired walkthroughs of the P2 and P3 interfaces by four members of the author team who did not participate in the design or implementation of Relic or the two interfaces. Participants were informed in advance that their interaction experience and feedback would be used for research purposes and reported as part of this study. These walkthroughs are reported as formative internal observations of the implemented interfaces.

Reproducibility Statement

The appendices document the experimental conditions, frozen workload construction and qualification, model and decoding configurations, random seeds, run protocol, prompts, protocol representations, metric definitions, and statistical procedures used in the reported experiments. They additionally specify evaluator isolation, artifact and configuration identities, checkpointing and replay, retry and recovery policies, and condition-conformance checks. The CooperBench and ProgramBench extensions are documented separately with their fixed evaluation sets, system configurations, and scoring procedures. These materials are intended to support reproduction and inspection of the reported experimental settings, comparisons, and analyses.

References

Ackerman, Mark S. 2000. “The Intellectual Challenge of CSCW: The Gap Between Social Requirements and Technical Feasibility.” Human–Computer Interaction 15 (2–3): 179–203. https://doi.org/10.1207/S15327051HCI1523_5.
Bakouch, Elie, and Prime Intellect. 2026. Measuring Autonomous AI Research. Prime Intellect Blog. https://www.primeintellect.ai/blog/measuring-autonomous-research.
CooperBench. 2026a. CooperBench Coordination Study: Agent Trajectories. Hugging Face Dataset. https://huggingface.co/datasets/CooperBench/team-trajectories.
CooperBench. 2026b. CooperBench: Official Implementation. GitHub repository. https://github.com/cooperbench/CooperBench.
Du, Hongyi, Jiaqi Su, Jisen Li, et al. 2026. “ProtocolBench: Which LLM MultiAgent Protocol to Choose?” Proceedings of the 43rd International Conference on Machine Learning. https://arxiv.org/abs/2510.17149.
Feldman, Martha S., and Brian T. Pentland. 2003. “Reconceptualizing Organizational Routines as a Source of Flexibility and Change.” Administrative Science Quarterly 48 (1): 94–118. https://doi.org/10.2307/3556620.
Hao, Zhezheng, Tianfu Wang, Huanshuo Dong, et al. 2026. Evolve as a Team: Collaborative Self-Evolution for LLM-Based Multi-Agent Systems. https://arxiv.org/abs/2605.29790.
Hong, Sirui, Mingchen Zhuge, Jiaqi Chen, et al. 2023. “MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework.” arXiv Preprint arXiv:2308.00352. https://arxiv.org/abs/2308.00352.
Khatua, Arpandeep, Hao Zhu, Peter Tran, et al. 2026. CooperBench: Why Coding Agents Cannot Be Your Teammates yet. https://arxiv.org/abs/2601.13295.
Lee, Charlotte P. 2007. “Boundary Negotiating Artifacts: Unbinding the Routine of Boundary Objects and Embracing Chaos in Collaborative Work.” Computer Supported Cooperative Work 16 (3): 307–39. https://doi.org/10.1007/s10606-007-9044-5.
Li, Guohao, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. 2023. “CAMEL: Communicative Agents for ‘Mind’ Exploration of Large Language Model Society.” arXiv Preprint arXiv:2303.17760. https://arxiv.org/abs/2303.17760.
Lyu, Wenhan, Yue Xiao, Yixuan Zhang, and Yifan Sun. 2026. Self-Organizing Multi-Agent Systems for Continuous Software Development. https://arxiv.org/abs/2603.25928v1.
Nelson, Richard R., and Sidney G. Winter. 1982. An Evolutionary Theory of Economic Change. Harvard University Press.
Qian, Chen, Wei Liu, Hongzhang Liu, et al. 2023. “ChatDev: Communicative Agents for Software Development.” arXiv Preprint arXiv:2307.07924. https://arxiv.org/abs/2307.07924.
Rajasekaran, Prithvi. 2026. Harness Design for Long-Running Application Development. Anthropic Engineering. https://www.anthropic.com/engineering/harness-design-long-running-apps.
Schmidt, Kjeld, and Liam Bannon. 1992. “Taking CSCW Seriously: Supporting Articulation Work.” Computer Supported Cooperative Work 1 (1–2): 7–40. https://doi.org/10.1007/BF00752449.
Shinn, Noah, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. “Reflexion: Language Agents with Verbal Reinforcement Learning.” Advances in Neural Information Processing Systems 36. https://papers.neurips.cc/paper_files/paper/2023/hash/1b44b878bb782e6954cd888628510e90-Abstract-Conference.html.
Teece, David J., Gary Pisano, and Amy Shuen. 1997. “Dynamic Capabilities and Strategic Management.” Strategic Management Journal 18 (7): 509–33. https://doi.org/10.1002/(SICI)1097-0266(199708)18:7<509::AID-SMJ882>3.0.CO;2-Z.
Wang, Guanzhi, Yuqi Xie, Yunfan Jiang, et al. 2023. Voyager: An Open-Ended Embodied Agent with Large Language Models. https://arxiv.org/abs/2305.16291.
Wu, Qingyun, Gagan Bansal, Jieyu Zhang, et al. 2023. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. https://arxiv.org/abs/2308.08155.
Yang, John, Carlos E. Jimenez, Alexander Wettig, et al. 2024. “SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.” The Thirty-Eighth Annual Conference on Neural Information Processing Systems. https://arxiv.org/abs/2405.15793.
Yang, John, Kilian Lieret, Jeffrey Ma, et al. 2026. ProgramBench: Can Language Models Rebuild Programs from Scratch? https://arxiv.org/abs/2605.03546.
Young, Justin. 2025. Effective Harnesses for Long-Running Agents. Anthropic Engineering. https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents.
Yu, Zhengxu, Yu Fu, Zhiyuan He, et al. 2026. From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company. https://arxiv.org/abs/2604.22446.
Zhang, Guibin, Muxin Fu, Guancheng Wan, Miao Yu, Kun Wang, and Shuicheng Yan. 2025. G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems. https://arxiv.org/abs/2506.07398.
Zhao, Andrew, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, and Gao Huang. 2023. ExpeL: LLM Agents Are Experiential Learners. https://arxiv.org/abs/2308.10144.
Zhu, Kunlun, Hongyi Du, Zhaochen Hong, et al. 2025. “MultiAgentBench: Evaluating the Collaboration and Competition of LLM Agents.” arXiv Preprint arXiv:2503.01935. https://arxiv.org/abs/2503.01935.

A Extended Method and Formal Specification

This section specifies organizational state, action selection, governance, execution, and capability formation.

A.1 State and Member Information

Let HtH_t denote the complete recorded history of a run. For member ii, let Ii,tacc⊆HtI^{\mathrm{acc}}_{i,t}\subseteq H_t be the information the member is allowed to access, and let vi,t⊆Ii,taccv_{i,t}\subseteq I^{\mathrm{acc}}_{i,t} be the bounded context actually read or retrieved at decision time. The runtime maintains vi,t⊆Ii,tacc⊆Ht.v_{i,t}\subseteq I^{\mathrm{acc}}_{i,t}\subseteq H_t. Private branches, unread messages, private workspaces, and evaluator-only assets therefore need not appear in a member’s context even though they exist in the recorded run.

A.1.1 Member-Local State

Member-local state contains the information and learned quantities carried by one member: functional role, decision prior, skills, current work, private workspace, private memory, local commitments, and bounded workload state. These quantities can affect that member’s later decisions without becoming organization-owned state.

Table 5. Member-local state.
Member-local state used in the experiments
CategoryRole in the system
Identity and functionRole, specialty, and stable functional decision prior.
Skill and standingSkill, reputation, and authority variables used by the structured selector where enabled.
Current workActive tasks, work status, and local scheduling state.
Private workspaceMember-owned notes, files, drafts, branches, and sandbox outputs.
Private memorySalient previously observed events retained for later context.
CommitmentsMember-specific promises, requests, and unresolved obligations.

A.1.2 Organizational State

We write organizational state as Ot=(Wt,Qt,Rt),O_t=(W_t,Q_t,R_t), where WtW_t contains shared work objects, QtQ_t contains responsibility and decision-right mappings, and RtR_t contains adopted organizational protocols. The defining distinction is operational: a shared document belongs to WtW_t; a rule belongs to RtR_t when it has an active runtime binding that can affect later work.

Table 6. Organizational state.
Organizational state
PartContentsOperational role
W_tTasks, issues, shared documents, repository artifacts, boards, messagesPersistent shared work and evidence.
Q_tTask ownership, reviewer/approver responsibility, authority relationsRoutes work and assigns organizational responsibility.
R_tAdopted protocol specifications, lifecycle state, use/violation/enforcement recordsStores learned ways of working that can govern later members.

A.1.3 Execution Records

Evaluator-side provenance records actions, work-object transitions, governance events, and decision traces for measurement and replay. These records preserve the sequence connecting proposal, adoption, later use, and execution consequences. Member decisions continue to use the bounded perception and retrieval interface in Appendix A.1; the complete recorded history supports researcher-side reconstruction of the same trajectory.

A.2 Persistent Organizations and Organizational Learning

A persistent organization is a multi-member system in which learned rules and cross-member responsibilities are stored outside member-local histories and can remain operative when members are replaced. We represent the organization at time tt as At=(Nt,{mi,t}i∈Nt,Ot,(C,Dθ,E),Γ),\mathcal{A}_t = \bigl(N_t,\{m_{i,t}\}_{i\in N_t},O_t,(\mathcal{C},D_\theta,E),\Gamma\bigr), where NtN_t is the roster, mi,tm_{i,t} is member-local state, OtO_t is organizational state, (C,Dθ,E)(\mathcal{C},D_\theta,E) is the fixed candidate, decision, and execution substrate, and Γ\Gamma is the governance process that introduces or revises persistent mechanisms.

A.2.1 Member Learning and Organizational Learning

Let It=(Qt,Rt)I_t=(Q_t,R_t) denote the institutional component of OtO_t, containing responsibility and decision-right mappings and adopted protocols. The experiment distinguishes three update paths: Wt+1=Fwork(Wt,at),It+1=Finstitution(It,Wt,at,gt),mi,t+1=Fagent(mi,t,oi,t,at,rt).\begin{aligned} W_{t+1} &= F_{\mathrm{work}}(W_t,a_t),\\ I_{t+1} &= F_{\mathrm{institution}}(I_t,W_t,a_t,g_t),\\ m_{i,t+1} &= F_{\mathrm{agent}}(m_{i,t},o_{i,t},a_t,r_t). \end{aligned}(2) Ordinary work changes shared artifacts; member learning changes a member’s local state; organizational learning introduces or revises governed mechanisms whose effects can extend to members who did not create them.

A.2.2 Persistence Across Member Replacement

Member replacement reinitializes private memory, private workspace and sandbox state, message read/acknowledgement state, commitments, reflections, wishes, and accumulated member-local experience. The functional role slots are reinstantiated, and shared work can persist across the replacement. In the reported transfer experiment, a freshly initialized target receives the specified protocol package and responsibility mappings. The target initialization and transferred contents are specified in Appendix E.2.1.

A.3 Structured Decision Layer and Execution

At each decision, the system constructs a feasible candidate set Ci,t\mathcal{C}_{i,t} from member-visible work state, role constraints, target validity, and the common action surface. Candidate construction is shared by the compared conditions; the action-selection mechanism differs by condition.

A.3.1 Structured Decision Layer (SDL)

B2 and B3 score candidates with the same 32-dimensional, rule-computed feature representation. The feature groups are shown in Table 7.

Table 7. SDL feature groups used by B2 and B3.
SDL candidate-feature dimensions
Feature groupDimensions
Task progress and priority7
Skill and role fit4
Social coordination and trust8
Repository health and evidence quality6
Protocols and institutional memory5
Governance-proposal signals2
Total32

For candidate aa, wk(i,t)=bk+∑jpij(t)βjk,sbase(a)=∑k=132wk(i,t)ϕk(a,xi,t),w_k(i,t)=b_k+\sum_j p_{ij}(t)\beta_{jk},\qquad s_{\mathrm{base}}(a)=\sum_{k=1}^{32}w_k(i,t)\phi_k(a,x_{i,t}), and u(a)=sbase(a)+A(a)+R(a)+C(a)+∑gGg(a).u(a)=s_{\mathrm{base}}(a)+A(a)+R(a)+C(a)+\sum_gG_g(a). The stochastic branch samples from P(a)=exp⁡((u(a)+ϵa)/0.6)∑a′exp⁡((u(a′)+ϵa′)/0.6),ϵa∼Uniform(−0.05,0.05),P(a)=\frac{\exp((u(a)+\epsilon_a)/0.6)} {\sum_{a'}\exp((u(a')+\epsilon_{a'})/0.6)}, \qquad \epsilon_a\sim\mathrm{Uniform}(-0.05,0.05), using namespaced seeded randomness. The scoring parameterization is fixed before the reported runs and shared by B2 and B3.

Table 8. Complete base-weight vector, grouped for readability.
SDL base-weight coefficients
GroupFeatureBase weight
Task / progressprogress gain; deadline urgency; task priority; blocker resolution; dependency unlock; demo relevance; customer relevance+0.55,+0.30,+0.30,+0.35,+0.20,+0.30,+0.30
Skill fitskill match; role affinity; learning gain; low-skill failure risk+0.40,+0.30,+0.10,-0.30
Social / coordinationcoordination; clarity; trust gain/risk; conflict risk; visibility; reputation gain/risk+0.30,+0.20,+0.25,-0.25,-0.25,+0.20,+0.20,-0.20
Repository / evidencerepository health; technical debt; review quality; reproducibility; untracked-result risk; claim evidence+0.30,-0.30,+0.25,+0.25,-0.20,+0.25
Protocolcreation potential; use potential; violation risk; enforcement gain; institutional memory+0.20,+0.20,-0.40,+0.20,+0.20
Governanceproposal endorsement; proposal skepticism+0.45,+0.45
Table 9. Post-score terms used by the SDL.
SDL post-score terms
Softmax temperature0.6
Seeded jitterUniform(-0.05,+0.05)
Authority bonus0.07× domain authority
Reputation bonusclipped at 0.10
Coding affinitybounded role–action bonus
Table 10. Principal soft-guard terms outside the core dot product.
Principal soft-guard penalties
Repetition-0.18
Artifact churn-0.12
Unresolved review-0.15
Missing new information-0.20
Opportunity cost-0.12
Shipping/off-task-0.50
Unshipped release idle-1.20

Algorithm A1. SDL-based collective decision (one member, one step)

Construct visible feasible candidates Ci,t\mathcal{C}_{i,t}. Apply shared feasibility, cooldown, and pool-size guards. For B0/B1, the model chooses from the candidate set. For B2/B3, extract the 32 features, compute u(a)u(a) for each candidate, and select by deterministic maximum or seeded softmax. Record the choice trace for evaluation, then execute the selected action through the common execution substrate.

A.3.2 Execution Operator

The common execution layer validates the selected action, applies its work-state transition, accounts for costs, and records the resulting events. In standard B3, adopted protocols supply decision features and responsibility signals to the selector, followed by protocol-linked checks and records after the action handler. SDL utilities determine action priorities and selection probabilities; transition validation and execution checks determine the resulting work-object disposition. Shared repository requirements, including review and current-CI evidence, are enforced by the common environment. The action registry and execution handlers remain fixed while adopted protocol specifications and role mappings evolve.

A.4 Capability Formation and Governance

Relic converts visible recurring work friction into proposals. A proposal names the problem, supporting evidence, affected work, required actions, and the intended organizational response. Validation checks grounding, supported actions, scope, duplication, and governance requirements before a proposal can be adopted.

A.4.1 Proposal and Protocol Representation

Table 11. Proposal and adopted-protocol fields.
Proposal and executable-protocol representation
GroupFields
Identityproposal type, title, proposer
Evidencesource observations/wishes, target problem, proposed response
Feasibilityrequired actions and capabilities
Governancereview state, required approvals, support/opposition, adoption score
Lifecycleamendment target, revision state, retirement state
Protocol contenttrigger, scope, responsible roles, required steps/evidence, affected actions/artifacts, violation/exception rules, enforcement response, success metric
Accountinguse, violation, enforcement, revision, and last-use records

A.4.2 Governance, Adoption, and Runtime Binding

The reported main-study configuration uses semi-automatic governance: members provide approvals first, while eligible proposals that remain pending for at least 24 steps can receive bounded deadlock recovery. Adoption additionally requires at least two supporters, at least two distinct approvers, a minimum review interval of three steps, and an adoption score of at least 0.60.6.

Table 12. Approval modes supported by the governance layer.
Approval-mode semantics
AutoEligible proposals can receive required approvals automatically.
AgentRequired approvals must come from member actions.
Semi-autoMember approval is primary; bounded deadlock recovery can complete missing approvals after the waiting period.

After adoption, a structured protocol specification is registered as organization-owned state. Its trigger, responsibilities, evidence obligations, and affected actions become available to the decision and execution layers. Amendments follow the same governed path; retirement disables future use while retaining lifecycle history.

Algorithm A2. Governed organizational-state update

From visible repeated friction, create a proposal. Validate its grounding and runtime compatibility. Route it for governance. When support, approval, review-time, and adoption-score requirements are satisfied, compile the proposal into a structured protocol and register it as organizational state. Later matching actions can generate use, violation, or enforcement records. Amendments repeat the same governance path; retired protocols cease governing future work while their history remains auditable.

A.4.3 Operational Capability-Formation Criterion

A registered autonomous protocol is counted as formed when it has a valid proposal and adoption record, at least two qualifying post-adoption uses, those uses span at least two recorded organizational units and two distinct steps, and the span from adoption to the last qualifying use is at least 48 steps. Strong evidence additionally links later third-party execution or enforcement to a governed-object state change and independent evaluator evidence.

Algorithm A3. Offline capability detection (read-only)

For each adopted autonomous rule, collect qualifying post-adoption uses. Mark weak formation when the use-count, organizational-unit, distinct-time, and 48-step persistence conditions all hold. Mark strong evidence when weak formation is additionally linked to later third-party execution or enforcement, a governed-object state change, and independent evaluator evidence. The detector reads the recorded ledger and does not modify the run.

A.5 Core Experimental Invariants

Table 13 summarizes the implementation invariants linking member-visible execution, governed updates, controlled comparisons, and measurement. Configuration and access checks are detailed in Appendix D.6.

Table 13. Core invariants enforced by the experimental design.
Experimental invariants
Member information interfaceMember context is built from permitted information that has been read or rendered; evaluator-private assets are held in the evaluator interface.
Cross-member privacySharing explicitly controls the visibility of private branches, workspaces, messages, and sandbox outputs to other members.
Governed institutional updatesProtocol adoption follows proposal validation and the configured review and approval path.
Registered action surfaceState mutations occur through registered actions and validated transitions.
Matched B2/B3 substrateB2 and B3 share the structured selector, candidate construction, member-learning path, task information, and execution substrate; B3 enables institutionalization.
Fresh-member resetTransfer reinitializes member-local state and applies the designated organizational package and role mappings.
Read-only measurementFormation detection reads recorded ledgers; evaluators score exported candidate state outside the organizational rollout.

B Execution-Grounded Organizational Environment

This section summarizes the common software-production environment used by all main-study conditions. The environment provides persistent shared work, partial observability, executable work transitions, and a frozen evaluator. Condition-specific organizational learning is layered on top of this common substrate.

B.1 Environment and Work State

Members act on persistent software-production objects: tasks and issues, working-tree artifacts, branches, commits, pull requests, CI records, reviews, shared documents, messages, and release state. Product progress is measured from concrete state transitions.

B.1.1 Main-Study Components

Table 14. Environment components relevant to the reported main study.
Common environment and condition-specific activation
ComponentAvailableMain-study role
Repository, CI, review, merge, releaseyesCommon execution-grounded work substrate.
Tasks, issues, shared boardyesCommon ownership and responsibility substrate.
Channels and direct messagesyesCommon information-friction substrate.
Private workspaces and sandboxesyesCommon partial-observability substrate.
Structured decision layeryesUsed for action selection in B2/B3.
Member learningyesEnabled in B2/B3.
Protocol proposal, governance, and executionyesInstitutionalization enabled in B3.
Frozen evaluatoryesScores candidate mainline state outside member context.
Token accountingyesMeasures realized model usage.

B.2 Shared Work, Visibility, and Communication

The environment separates private, shared, and evaluator-only state. Local edits remain private until committed and exposed through the repository workflow; sandbox outputs remain private until explicitly shared; a shared document can be visible before its full content is opened; and messages become part of a member’s context only after the member reads them. Evaluator-only tests, held-out requirements, and references never enter member-visible state.

B.2.1 Read and Retrieval Semantics

Visibility, delivery, reading, acknowledgement, and memory are distinct. Inbox triage admits at most 25 unread perceivable messages per step. Search operates over frozen internal corpora, repository state, member-local sandbox state, and fixed external-snapshot corpora; no live network retrieval is used in the reported environment.

B.3 Action Surface and Execution Validation

The common action surface covers task work, repository operations, verification, communication, meetings, search, documents, artifacts, governance, and release operations. Feasible actions are generated from the member’s visible state and role. Selection is followed by transition-specific validation, so an offered action can still fail when its target or precondition has changed.

B.3.1 Invalid Actions and Failures

Table 15. High-level invalid-action handling used in the experiments.
Execution outcomes
Failure classWhere resolvedRecorded effect
Invisible or invalid targetcandidate/transition validationNo product mutation; decision is rejected.
Invalid work transitionaction-specific validationFailed execution with recorded reason.
Missing role or approvalaction/governance validationTransition refused.
Malformed model outputstructured-output validationNo action mutation; model usage remains counted.
Unavailable actioncandidate constructionAction is absent from the feasible set.

B.4 Software-Production Lifecycle and Verification

Repository-based workloads begin from a frozen initial repository and specify repairs or feature extensions. Construction workloads begin from an unimplemented starter skeleton plus a frozen specification and public checks. In both cases, the delivery chain is edit →\rightarrow commit →\rightarrow pull request →\rightarrow review/CI →\rightarrow merge, and reported product metrics are evaluated on mainline.

Verification is commit-bound. A CI result records the tested pull-request head and the mainline state against which it was evaluated. Pushing a new commit invalidates earlier evidence for that head, and mainline movement can make a previous result stale before merge. These checks are common environment properties shared by all conditions; Relic can learn organizational rules that route responsibility and evidence around them.

B.5 Time, Scheduling, and Resource Accounting

The main-study horizon is 336 scheduling steps with checkpoints every 24 steps. A scheduling step denotes one simulation update; elapsed execution time is recorded separately in seconds. Members take at most one primary action per step, while background work can complete asynchronously. Runtime records retain the field name tick for the scheduling index.

Realized model usage is accumulated from provider receipts, including retry calls. Tool and container compute is unmetered and remains separate from the provider-token totals. Checkpoint resume loads the saved organization and continues under the same run identity, preserving the relation between the trajectory, its scheduling position, and its accumulated resource use.

B.6 Organizational Runtime and Protocol Activation

The shared work layer stores operational artifacts; responsibility mappings identify owners, reviewers, and approvers; adopted protocols add triggers, required evidence, affected actions, responsible functions, exception rules, and lifecycle state. In standard B3, learned rules can reprioritize actions, route responsibility, and participate in protocol-linked execution checks. Amendments follow the same governance path as initial adoption, and retirement stops future application while retaining the recorded history.

B.7 Logging, Checkpointing, and Replay

Each executed action records its actor, action type, success or failure, affected work objects, costs, and resulting events. Governance transitions add proposal, adoption, amendment, and retirement records; protocol execution adds the corresponding use and enforcement events. Checkpoints serialize organizational state, member state, and scheduling state for run recovery.

Case-study timelines are reconstructed from these ledgers and checkpoints. Researcher-facing inspection joins the complete recorded state and event history; member-facing decisions use the bounded perception interface. This separation preserves both the local information used to choose an action and the shared-state consequences used to analyze it.

C Benchmark Curation

This section documents the ten frozen software-production workloads used in the 360-run headline endpoint study: three models, three seeds, and four conditions per workload. W01–W10 are stable identifiers in the frozen analysis PACK_ORDER. Table 16 gives their task names, version transitions, and original pack IDs; the same mapping applies to all workload-level results, transfer references, and the selected case. Construction and pre-run qualification are specified below.

Development tasks.

Relic was developed and debugged on repositories outside W01–W10 before the reported evaluation. Development records and the 360-run evaluation matrix have separate run identities. Within B3 evaluation runs, online protocol formation and revision remain active components of the method.

C.1 Evaluation Scope and Workload Taxonomy

C.1.1 Frozen-Version Repair and Feature Development

W06–W10 use existing upstream projects with both an initial version and a final reference version frozen during benchmark construction. The complete initial repository is the starter; the final repository is the reference implementation. The exact version pairs appear in Table 16. Feature changes and issue requirements between these versions define the target behaviors, covering both bug fixes and new functionality. Behavioral test cases and verification logic are written for those requirements and executed against both versions. Every retained scored case must fail on the untouched initial version and pass on the frozen final version. The resulting requirements, cases, verifier, and version pair are fixed before organizational evaluation.

Members receive the starter, designated public requirements, and public checks. The final implementation, hidden acceptance tests, and held-out requirements remain evaluator-side during rollout. A verified repair or feature completion requires a retained baseline-failing check to pass on the candidate, using baseline_status == failed and candidate_status == passed. The frozen analysis retains the label repair for the W06–W10 family; this repository-based family includes both defect repair and feature development.

C.1.2 Cross-Cutting Feature Extension

Cross-cutting requirements modify public interfaces and dependencies across modules. The interface-repair affordance inserts an edit candidate for each importer when an earlier edit removes a public symbol. The parameter-threading affordance similarly supplies edit candidates when a signature gains a parameter that is not forwarded. These candidates expose the dependent implementation work through the common action surface.

The frozen manifests distinguish complete repository snapshots from construction skeletons, while component maps identify the module boundaries touched by each requirement. Cross-cutting extension demand is recorded within the repository-based and construction workload families.

C.1.3 Zero-to-One Construction

W01–W05 are zero-to-one construction workloads grounded in the functionality of existing software. For each workload, the author manually selected an existing project as a functional reference and specified the capabilities of a new, analogous program. The authoring LLM received this functional brief without the original project’s source code or existing implementation. It first wrote technical documentation and an explicit feature specification, then implemented a new reference program, and subsequently wrote feature test cases and a verifier from the documented features and that completed implementation.

An unimplemented starter skeleton supplies the new program’s package structure and signatures. Every retained scored case is checked against both this starter and the completed reference, requiring failure on the former and success on the latter. Requirements, starter, reference, cases, and verifier are frozen before any evaluated organization begins work. The evaluated members receive the specification, starter, and designated public checks; they do not receive the original project’s implementation, the newly authored reference implementation, or evaluator-only test content.

Contracts group related behavioral cases; a complete contract requires every constituent case to pass. Mini Blobstore (W01) has five progressive contracts with seven cases each. Its qualified starter passes 0/35 cases and 0/5 contracts; its reference passes 35/35 and 5/5. All ten workloads satisfy the per-case qualification in Appendix C.5.3.

Dependencies between contracts create recurring implementation, verification, and integration demands. The workload supplies these dependencies without prescribing a particular division of labor.

C.2 Workload Sourcing and Eligibility

C.2.1 Frozen Sources and Specifications

Table 16. Frozen main-study workload mapping.
Frozen main-study workload mapping
IDBenchmark workloadFrozen pack ID
W01–W05: function-grounded zero-to-one construction
W01Mini Blobstoremini_blobstore_v1
W02Traffic Watchtraffic_watch_v1
W03TG Automationtg_automation_v1
W04PDF Reformatterpdf_reformatter_v1
W05FastAPI Dashboardfastapi_dashboard_v1
W06–W10: upstream version-transition tasks
W06Boltons 24.0.0 → 26.1.0boltons_v2400_to_v2610
W07Celery 5.6.0 → 5.6.3celery_v560_to_v563
W08Soup Sieve 2.6 → 2.9.1soupsieve_v26_to_v291
W09cattrs 25.1.0 → 26.1.0cattrs_v2510_to_v2610
W10Tenacity 8.2.3 → 9.1.4tenacity_v823_to_v914

W01–W10 identify the same workloads in task manifests, result tables, transfer records, and case-study references. Each frozen manifest links the starter, requirements, evaluator assets, reference, qualification result, and freeze record. Original pack IDs and upstream version pairs are retained in Table 16, providing the join from paper-level workload identifiers to the frozen source artifacts.

C.2.2 Inclusion Criteria

A pack is usable only when three conditions hold and tools/preflight_pack_environment.py --score reports all three:

  1. the environment works – every seeded issue names a file an agent can reach and edit, and the merge gate can actually run and refuse;

  2. the verifier works – every retained scored case fails on the untouched starter and passes on the fixed reference, verified case by case;

  3. the gate says something – the agent-visible acceptance gate is red on the untouched starter, so it can distinguish work from no work.

Additional structural requirements are enforced by tools/check_pack_health.py: every issue must resolve to an editable artifact; held-out requirements must be absent from the agent-visible stream; at least _MIN_PUBLIC_TESTS = 3 public tests must exist; and no editable file may exceed _MAX_EDITABLE_BYTES = 200,000. Dependencies must be pinned in-tree, the licence must permit redistribution, and no task may require a live network.

Each retained pack includes agent-visible acceptance checks that are red on the starter and green on the reference. The retained public-gate coverage is 12/12 for W06, 8/12 for W07, 6/9 for W08, 7/12 for W09, and 9/12 for W10.

C.2.3 Exclusion Criteria

Packs are excluded for unstable evaluation, manual scoring, incompatible licensing, already-satisfied starter behavior, unavoidable future-solution leakage, irreproducible dependencies, or absent organizational demand.

C.3 Freezing and Manifest Construction

C.3.1 Starter Artifacts

For W06–W10 the starter is the complete frozen initial repository version, including its dependency lock files (uv.lock, pdm.lock, Cargo.lock as the project uses); the final version is reserved as the reference. For W01–W05 the starter is an unimplemented skeleton of the new program specified by the authoring LLM from a human-selected functional brief. It supplies package structure and signatures plus public contract tests, without the original project’s implementation or the newly authored reference code. Both kinds declare their entry points and their agent-runnable public-test command in the manifest (entrypoints.cli, entrypoints.smoke, public_tests.command, public_tests.dependencies with exact pins).

The starter tree is identified by starter_repo_digest. Frozen dependency fixtures are documented in frozen_fixture_exemptions.json, with schema frozen_pack_fixture_exemptions_v1. Its 20 entries record the fixture path, pack and tree, scanner rule, SHA-256, byte count, source reference, and retention reason. The allowlist preserves the byte-identical fixtures required by the pinned dependencies.

C.3.2 Evaluator Artifacts

Each pack contains tests/hidden/, including the hidden suite and specs.json scoring-unit declaration; issues/public/ for agent-visible requirements; issues/heldout/ for evaluator-only requirements; and reference_repo/ for the reference implementation. The manifest freezes scoring-unit counts, qualification outcomes, and the entry points used by the evaluator.

Hidden assets are loaded into an in-process vault keyed by a 32-byte random token. World serialization retains an oss_evaluator_vault_binding_v1 token binding; the vault resolves the private content at evaluation time using constant-time comparison. run_oss_hidden_tests and materialize_and_run_oss_hidden_tests return aggregate fields such as passed, failed, total, pass_rate, issue_fix, and issue_fix_rate.

Qualification and candidate scoring use build_time_machine_evaluation_plan and evaluate_time_machine_candidate. Per-contract admission is resolved through oracle_blocking_reasons. These entry points connect the frozen cases, baseline status, and evaluated candidate state to the scoring records described in Appendix F.1.2.

C.3.3 Containers and Dependencies

The container evaluator uses a digest-pinned base image and a per-pack layer containing pinned test dependencies. The formal container path accepts docker or apptainer, requires trust_level = untrusted and network_enabled = False, and validates an image reference of the form sha256:[0-9a-f]{64}. The execution platform is linux/amd64. The repository build script creates the image, and ORG_EVALUATOR_CONTAINER_IMAGE supplies its concrete identifier at the execution site. A formal_digest_pinned_image_required status records a failed image-policy check.

Network isolation combines process-level and test-wrapper controls. HTTP_PROXY, HTTPS_PROXY, and ALL_PROXY point to the closed loopback endpoint http://127.0.0.1:9; the hidden-test wrapper restricts socket.socket.connect to loopback. The search system operates on its internal and frozen corpora through the separate interface in Appendix B.2.1.

The host development path evaluates an exported candidate tree in a separate process. Execution-policy and evaluator receipts identify the configuration used by each run, together with the candidate, test, and evaluator digests. Container execution and host-process execution therefore have explicit, recorded configurations for reproducing the corresponding result.

C.3.4 Versioning and Integrity Receipts

The manifest schema org_env_oss_time_machine_pack_v2 records starter_repo_digest, reference_repo_digest, hidden_suite_hash, evaluator_environment_hash, qualification_hash, qualification_result_hash, result_hash, runtime_test_status, agent_filesystem_isolation_status, oracle_reachability, and component_map, alongside the entry-point and test declarations. Run records carry the artifact-identity fields that associate an evaluated outcome with its frozen pack.

evaluator_environment_hash is computed from the evaluator source. A scoring-code change produces a new evaluator identity, while the recorded results retain the identity of the evaluator that produced them. The verification chain joins the dataset manifest, starter, reference, hidden suite, evaluator, tool surface, condition, event graph, and run manifest. Checkpoints additionally carry payload and sidecar digests. The recorded digest values and their frozen manifests support artifact matching during qualification, aggregation, and resume.

C.4 Task and Evaluator Construction

C.4.1 Authoring Provenance and Freeze Order

For W01–W05, authoring proceeds from the human-selected functional brief to technical documentation and feature specification, a new reference implementation, and then feature tests and verification logic. The original authoring trajectories are retained with the constructed artifacts and endpoint-qualification records. These records document how the specification, reference, and retained scoring cases were produced.

For W06–W10, the frozen initial/final repository pair anchors the feature and issue requirements and their acceptance checks. Both construction paths qualify each retained case on the starter and reference before B0–B3 rollouts. The resulting tests, references, and scoring configuration remain fixed during the evaluated organizational runs.

C.4.2 Agent-Visible Task Briefs

Briefs are rendered from frozen pack manifests. Repository-based briefs contain the product summary, public bug-fix and feature requirements, acceptance hints, entry points, and public-test command. Construction briefs contain the specification, progressive steps, skeleton layout, and public contract tests. Members also receive the artifacts and results shared during their run.

C.4.3 Exposed and Held-Out Behavioral Cases

An exposed unit corresponds to a requirement the organization can see: a seeded public issue, or a step of a published progressive contract. A held-out unit is a requirement in issues/heldout/ that is never shown to members. Its contracts are scored offline on final mainline states and saved checkpoints across all three models. In both cases the content of the scoring test is evaluator-only; what an exposed unit gives the organization is the requirement, plus a public acceptance check derived from it, not the hidden assertion.

The denominators depend on the scoring granularity declared by each pack, which is separate from its repair or construction label. In the frozen manifest records, contract-level scoring gives 16 units for W07, 13 for W08, 4 for W04, and 9 for W03. Other packs decompose hidden contracts into behavioural cases: W01 has 5 contracts scored in 35 cases and W02 has 9 contracts scored in 40 cases. Appendix F defines the rate aggregation for these scoring units.

C.4.4 Progressive and Complete Contracts

A complete contract requires every constituent behavioral case to pass; partial credit remains at case level. In Mini Blobstore (W01), the five progressive contracts contain seven cases each. Pre-run qualification gives 0/7 passing cases for the starter and 7/7 for the reference within every contract, yielding 0/5 and 5/5 complete contracts, respectively.

Agent outcomes are evaluated on mainline. Workspace results are reported separately (Appendix G.3.2).

C.5 Quality Control and Preflight

C.5.1 Starter and Reference Sanity

The preflight entry points are:

PYTHONPATH=. python tools/preflight_pack_environment.py --score <pack>
PYTHONPATH=. python tools/check_pack_health.py <pack>
PYTHONPATH=. python tools/print_evaluator_hashes.py <pack>

The first command executes the three qualification gates in Appendix C.2.2, including actual scoring of the starter and reference. The second checks issue-to-file routing, artifact resolution, oracle qualification, held-out visibility, public-suite size, and editable-file size. The third prints the evaluator-environment and qualification-plan hashes used as required batch-runner arguments.

Qualification receipts are written to manifest.yaml through runtime_test_status, qualification_hash, and qualification_result_hash, and to provenance/authoring_report.json. The launch preflight validates these receipts against the selected frozen workload.

C.5.2 Evaluator Determinism

Evaluator-source hashing supplies evaluator_environment_hash; the qualification payload supplies the plan hash. Scoring hashes the candidate tree before and after execution and raises an error on a detected mutation. The recorded tree, evaluator, and plan identities link each score to the exact evaluated state.

Malformed or erroring qualification contracts receive contract-level error statuses and are excluded by the qualification gate. CI and public-test infrastructure faults have the explicit statuses ci_infrastructure_error and public_tests_infrastructure_error. Test outcomes and execution faults consequently retain their own status fields through analysis.

C.5.3 Lower and Upper Bounds

Every retained scored case in W01–W10 is qualified against both the untouched starter and the fixed reference. The starter fails each retained case and the reference passes it, yielding 0% and 100% qualification scores on each retained acceptance suite. W01–W05 use the newly authored reference programs; W06–W10 use the frozen final upstream versions.

Qualification failure blocks admission with an explicit reason, including baseline_not_failed:<test_id>:<status> and reference_not_passed:<test_id>:<status>. Test source, references, and scoring rules are frozen before organizational rollout and versioned independently of release packaging.

C.5.4 Leakage and Identity Checks

Four automatic checks validate the member-facing interfaces: artifact inspection checks hidden and reference content, search tests inspect retrievable results, perception tests inspect rendered packets, and snapshot tests verify the aggregate-count representation. The formal world also validates evaluator-vault bindings at initialization.

A post-run future-recall check compares identifiers written by agents with identifiers found only in evaluator-side future versions. Together with the runtime access checks, these tests connect the information invariants in Table 13 to concrete artifact, search, perception, snapshot, and output records.

C.6 Completed Main Study and Evaluation Coverage

C.6.1 Main-Study Matrix

The main study contains 360 runs: ten workloads, three seeds, and B0–B3 under each of three models. GPT-5.6 Terra, Claude Opus 4.6, and DeepSeek-V4-Flash each contribute 120 runs. Endpoint evaluation and 24-step checkpoint curves cover the three-model study.

C.6.2 Analysis Populations

Table 17. Analysis populations.
AnalysisRunsDesign
Main endpoints and checkpoint curves360Three models; ten workloads; three seeds; B0–B3.
Terra+Opus auxiliary analyses240Two models; ten workloads; three seeds; B0–B3.
Autonomous-protocol census60Terra+Opus B3 runs.

C.6.3 Workload Manifest Table

Table 18. Scoring-unit inventory in the frozen W01–W10 order.
Frozen scoring-unit inventory
PackFamilySeededExposedHeld-outContractsCasesStatus
W01construction5350535main study + case
W02construction772940main study
W03construction77299main study
W04construction33144main study
W05construction33144main study
W06repair121241616main study
W07repair121241616main study
W08repair9941313main study
W09repair121241616main study
W10repair121241616main study

Scoring-unit identity check.

For W03, the manifest declares seven public issue units and two held-out issue units. Scored B2/B3 endpoint rows carry exposed_total = 7 and heldout_total = 2, matching the nine contract-level behavioral cases in Table 18. This field-level join ties the exposure split to the frozen task definition.

D Experimental Conditions and Run Protocol

D.1 Exact B0–B3 Definitions

Conditions are declared as a frozen dataclass with eleven fields and resolved by condition identifier. The reported configurations are b0_single_agent_founder, b1_persistent_role_org, b2_policy_conditioned_org, and b3_full_relic. Each run retains the resolved condition and its configuration fingerprint.

D.1.1 B0: Reflective Individual

B0 uses one member with action_selection_mode = llm_direct. The member is constructed from the canonical eight-member roster: skill coverage is the per-domain maximum, while decision priors, communication style, and work rhythm use roster means. It has a private workspace, sandbox, and memory with the shared salience threshold and capacity limit, plus the common registered actions, tools, and task-visible information.

Operational tasks, artifacts, budget records, and repository state persist throughout the run. institutionalization_enabled and capability_learning_enabled are false. Appraised observations enter private memory under the shared salience and capacity rules; inspected or retrieved material enters the bounded context for subsequent LLM-direct decisions.

D.1.2 B1: Long-Horizon Role-Based Multi-Agent System

B1 uses eight persistent members with fixed functional roles, a shared company workspace, channels, meetings, and private member memory. The roster and member-local state persist from t=0t=0 through t=336t=336; temporary_team = false retains them across episode and sprint boundaries. The model selects from the feasible candidate pool through action_selection_mode = llm_direct.

profile_conditioning_enabled, capability_learning_enabled, and institutionalization_enabled are false. Role-scoped affordances, shared artifacts, communication, and repository delivery remain available through the common execution substrate. Shared workspace and private member memory retain their respective visibility and storage interfaces.

D.1.3 B2: SDL-Based Long-Horizon Multi-Agent System

B2 uses the same eight-member team and work environment as B1, with SDL action selection (action_selection_mode = profile_policy), profile conditioning, and member-local skill, reputation, and authority updates.

Protocol proposal, adoption, compilation, and revision are disabled. The registered institutional actions are propose_protocol, amend_protocol, support_protocol, and follow_protocol. They may enter the candidate menu; with institutionalization disabled, their handlers return mechanism_ablation:institutionalization. The wish-to-proposal cadence, adoption, compilation, and registration path are skipped, while member-local learning, private memory, and shared artifacts continue to accumulate.

B1–B2 compares the combined addition of SDL, profile conditioning, and member learning. B3–B2 enables the protocol lifecycle on that shared backbone.

D.1.4 B3: Persistent Protocol-Governed Organization

B3 shares the B2 roster, model, tools, candidate construction, features, SDL parameterization, member learning, and execution handlers. Enabling institutionalization_enabled activates reflection-to-wish, proposal, validation, governance, and protocol compilation. Adopted protocols enter RtR_t and QtQ_t, supplying decision features, responsibility signals, and post-handler protocol checks and records. Members can amend and retire adopted rules through the same lifecycle.

D.2 Roster, Functions, and Decision Priors

D.2.1 Roster Size and Functional Specialties

The canonical roster is eight members, used in full by B1, B2 and B3, and aggregated into one member for B0. The functional specialties are a leadership and direction function, two implementation functions differing in speed-versus-care emphasis, a reliability and verification function, and four further functions covering review, documentation, customer signal and coordination.

D.2.2 Roles, Responsibilities, and Permissions

The board starts with unassigned tasks, and members allocate ownership during execution. Role-scoped affordances determine which functions receive edit and review candidates and which may approve releases; these affordances are identical across B1, B2, and B3. Release approval requires a lead role and the reliability function, capped by roster size for multi-member conditions. Object permissions follow Appendix B.2, and escalation combines the registered escalation action with the governance deadlock rule.

The declared-factor conformance test checks that all multi-member conditions use the same governance role topology. The comparison therefore retains the same functional approval structure while varying the declared mechanism switches.

D.2.3 Functional Decision Priors

A prior pi(t)p_i(t) is a member-specific map of named, real-valued profile and skill dimensions initialized from the canonical roster. It enters the 32-dimensional SDL representation in Appendix A.3.1. The reported assignment is aligned, with functional conditioning fixed before the runs and shared by B2 and B3.

Profile-backed inputs remain fixed, while skill-backed inputs evolve through the common member-learning path. The profile_conditioning_enabled switch activates these prior inputs in B2 and B3. The same roster initialization and profile assignments are used for the two configurations.

D.2.4 Member Initialization

Every member starts with: empty private memory; an empty personal workspace and an empty sandbox; no owned tasks; default continuous condition (attention 1.0, morale 0.6, trust_in_company 0.7, perceived_recognition 0.5, role_clarity 0.4, the remaining variables 0); the roster’s literal prior and skill maps; and eight-domain reputation and authority at their configured baselines. Task knowledge is whatever the brief and the starter tree contain (Appendix C.4.2); nothing about the hidden suite, the reference or the held-out requirements is present in any member’s initial state.

Randomization at initialization is limited to availability registration, which is seeded. Under fresh-member turnover, the same initialization procedure reconstructs member-local state for the target organization while preserving the functional role slots. A fresh member therefore starts with the same kind of local state as a member at t=0t=0 of an ordinary run (Appendix E.2.1).

D.3 Prompts and Information Controls

Relic separates selection of what to do from model calls used to carry out or reflect on a selected activity. B0 and B1 ask an LLM to choose one action from a constrained candidate menu. B2, B3, B3-text, Transfer Text, and Transfer Exec use SDL for what-action selection and issue no LLM action-selection prompt. Other model calls remain module-specific, including code/document editing, reflection, proposal drafting, and proposal evaluation.

Table 19. Prompt routing across conditions. ``cond.'' denotes a module call when that activity occurs.
Prompt stageB0B1B2B3B3-textTextExec
Shared grounding / identitycond.cond.cond.cond.cond.cond.cond.
What-action selectionLLMLLMSDLSDLSDLSDLSDL
Code/document editingcond.cond.cond.cond.cond.cond.cond.
Reflectioncond.cond.cond.cond.cond.cond.cond.
Separate LLM wish extraction———————
Wish → proposal———cond.cond.——
Proposal evaluation———cond.cond.——
Approval-specific LLM call———————
Readable rules in code-editor context———avail.avail.66

The wish stage maps and filters reflection-generated improvement_ideas. Proposal evaluation supplies scores and revisions; governance actions supply approval and adoption.

D.3.1 Shared System and Grounding Template

Cognitive modules compose the system message from a shared workload/company brief, global grounding rules, the module name, and the module-specific instruction. The member identity block—role mandate, enabled functional conditioning, and current memory/context—is prepended to the user message. The grounding rules require the model to use supplied objects and identifiers, select only available actions and valid targets where applicable, avoid inventing outcomes, and return the requested structured output.

D.3.2 B0/B1 Action-Selection Prompt

For B0/B1, the user message presents the member-visible action context and asks for one exact candidate:

Choose the next intentional action for {agent_id} at tick {tick}.
{rendered_action_context}
Return ONLY JSON matching the action schema; candidate_action and target must identify one exact entry from candidate_options.

The returned choice is validated against the constrained candidate pool before execution. B2/B3 and all text/transfer variants do not use this prompt because SDL performs the corresponding selection.

D.3.3 Code and Document Editing

After an edit action has been selected, the code editor receives the target file, current content, available public evidence, the edit goal, and any readable organizational rules applicable to the treatment. The editor returns anchored replacements rather than a whole-file rewrite. In B3/B3-text and both transfer treatments, readable rules enter the model through a dedicated block:

HOW THIS ORGANIZATION WORKS:
- {readable rule 1}
- {readable rule 2}
…

D.3.4 Formation Prompts: Reflection, Wishes, and Proposals

Reflection receives recent work, failures, relevant episodes, product context, memory, and existing protocol context and asks for structured self/team assessment, blockers, and concrete improvement_ideas. A protocol_need is used when recurring avoidable failure suggests a standing, checkable rule. The main formation path then maps these ideas into wishes; no separate LLM wish-extraction call is used.

A qualifying wish is drafted into a proposal with the visible wish record, available registered actions, and product context. The proposal instruction requires a concrete checkable requirement, names required approvals, and does not allow the model to assume adoption. Proposal evaluation separately returns feasibility, usefulness, risk, adoption score, blocking issues, and suggested revision. Recurring-pattern synthesis can also produce a candidate protocol; it remains a proposal until the governance path adopts it.

The exact prompt text, readable transfer rules, and wish-to-protocol measurement criteria are given in Appendix I.7.1.

D.3.5 Context Rendering

Rendering is section-ordered and budget-aware. Candidate menus, task boards, inboxes, pull requests, test results, and product context retain the information needed to act; narrative history/reflection sections use bounded rendering. Inbox triage and private-memory capacity provide additional shared limits across conditions.

D.4 Matched Experimental Controls

D.4.1 Model and Decoding Configuration

The main study uses GPT-5.6 Terra, Claude Opus 4.6, and DeepSeek-V4-Flash, with 120 runs per model. The DeepSeek runs use the provider model identifier deepseek-ai/DeepSeek-V4-Flash-0731. GPT-5.6 Terra response identifiers include gpt-5.6-terra-2026-07-09 and the unversioned family name. Internal transfer uses GPT-5.6 Terra; ProgramBench uses GPT-5.6 with xhigh reasoning effort. The in-situ binding ablation uses Claude Opus 4.6. The same-model CooperBench comparison uses Claude Opus 4.6 with high reasoning effort.

Client transport documentation specifies a chat_completions route, a 180-second request deadline, six retries, and exponential backoff starting at three seconds. Connect, read, write, and pool timeouts have separate configuration fields. Each run retains its provider/endpoint routing, wire API, output limits, decoding fields, and request policy in the client configuration record. Reasoning effort and temperature are separately identified fields.

A model-call failure propagates as a failed decision or enters the module-specific fallback path. llm_runtime_identity records observed_response_models as a counter of returned model strings, an endpoint hash, a routing-context fingerprint, and a chained response-ID digest. llm_runtime_fingerprint hashes that identity block. Requested model identifiers, observed response identifiers, and routing metadata thus remain directly inspectable for each recorded configuration.

D.4.2 Tools and Registered Actions

One action registry and wiring function install the same candidate mapper, feature extractor, policy implementation, and execution adapters across the conditions. Condition switches choose LLM-direct or SDL action selection and activate member learning and institutionalization. The run record’s tool_surface_fingerprint identifies the offered tool and action surface used by that configuration.

Protocol and tool creation pass through B3 governance. The four institutional actions are named in Appendix D.1.3; their registry presence and condition-dependent execution are checked independently of the common repository-delivery actions.

D.4.3 Task-Visible Information

The pack and seed fix the starter tree, public requirements or specification, public-test command, entry points, and issue timing. Public-test outcomes are computed on each condition’s candidate state. The analysis joins paired conditions through the manifest, starter, hidden-suite, and evaluator identities retained in the run and evaluator records.

The aggregation pipeline validates starter digests before pooling paired conditions. A revised public gate or evaluator produces a versioned asset identity, and the analysis maps each outcome to its designated frozen task configuration. This join preserves the correspondence among the task, candidate artifact, scoring procedure, and reported comparison.

D.4.4 Execution Horizon and Resource Ceilings

The main-study matrix uses a 336-step horizon, checkpoints every 24 steps, 168-step sprints, and disabled work rhythm. Token expenditure is measured per run, with no matched hard token, model-call, or primary-action ceiling. Company-balance accounting is soft: charging clamps the remaining balance and records violation flags, while action execution remains governed by the action and repository validity checks.

The runner also supports a company-wide resource-ledger configuration with max_llm_calls, max_llm_requested_tokens, max_llm_prompt_characters, max_primary_actions, and max_ticks. When enabled, its append-only usage and exhaustion signal operate at organization scope. The run configuration records the active horizon and resource policy; the reported main-study results use the scheduling-horizon setting above.

D.4.5 Seeds and Randomness

Each model–workload–condition combination uses seeds 1401, 2711, and 4013. The mechanism case uses seed 1401.

Randomness is namespaced from the run seed rather than drawn from one global stream: policy jitter and softmax sampling use Random(f"{seed}:{tick}:{agent}"); availability registration, sandbox timing, the simulated community and substrate controls each derive their own stream; and the shuffled-prior transform uses a prefixed string seed. Model calls use the configured provider sampling. Within each step, members are served in roster insertion order.

D.5 Run Protocol

D.5.1 Initialization and Preflight

Before launch, the pack passes the three qualification gates with --score. The evaluator and qualification-plan hashes are printed and supplied as required runner arguments. The code tree is copied into a run-specific frozen directory, and the public suite is executed on both starter and reference to verify red/green acceptance behavior. The runner then initializes organizational state from the selected manifest.

Evaluator assets are loaded into the token-keyed vault, with the binding stored in world state. When prompt auditing is enabled, hashing starts at the first intercepted call and the audit receipt records its call coverage. The initialized run therefore retains the code, task, evaluator, and information-interface identities used by its subsequent execution.

D.5.2 Online Execution

Each step: advance the clock; at a day boundary run the daily passes (growth where enabled, budget, structural detectors); process meetings and background jobs; then for each member in roster order, triage the inbox, build perception, construct and guard candidates, select, execute, appraise into memory, and write graph edges. Governance passes run on their own cadences (Algorithm A2). State is written to a checkpoint every 24 steps together with token totals, recorded at the same checkpoint.

D.5.3 Checkpoint and Endpoint Grading

Checkpoint scoring evaluates saved states offline with a 120-second per-contract timeout. Endpoint grading evaluates final mainline state. The reported endpoint and 24-step checkpoint series cover the three-model, 360-run study, with the tree type, scoring units, and metric eligibility attached to each evaluated series.

Each checkpoint couples the saved organization and candidate state with its token totals and evaluated outcomes. Figure 4 uses these recorded checkpoint steps; the cost sensitivity selects saved checkpoints according to Appendix G.7.3.

D.5.4 Stopping, Submission, and Resume

Runs end at the scheduling horizon. Interrupted execution resumes from the last saved checkpoint under the same run identity. Appendix I.6.2 specifies provider retries and recovery.

D.5.5 Run Matrix and Inventory

Table 20. Experimental accounting.
Experimental accounting
Evidence lineRunsDesign / coverageConfiguration
GPT-5.6 Terra main study12010 3 4Ten workloads, three seeds, B0–B3.
Claude Opus 4.6 main study12010 3 4Ten workloads, three seeds, B0–B3.
DeepSeek-V4-Flash main study12010 3 4Ten workloads, three seeds, B0–B3; included in the pooled endpoint table.
Internal transfer6030 Text + 30 ExecTen workloads × three seeds per arm; 30 GPT-5.6 Terra B2 main-study references reused.
In-situ binding ablation30 newB3-text; Claude Opus 4.6, seeds 1401/2711/4013Thirty matched B3 references reused from the Claude Opus 4.6 main study.
Selected caseincludedOne matched GPT-5.6 Terra main-study unitW01; GPT-5.6 Terra; seed 1401.
ProgramBenchseparateSame 25 tasks, two systemsExternal protocol plug-in; no SDL.
CooperBenchseparateFull 652-pair benchmark + 47-pair same-model subsetTwo-member Relic; full released-reference comparison plus same-model Solo/Peer control using the same broken-pair exclusion rule.
Human–agent extensionwalk\-through studyP2/P3 walkthroughsHuman-facing interface evaluation.

D.6 Condition-Conformance Tests

D.6.1 Mechanism-Isolation Tests

The declared-factor test enumerates the fields allowed to change at each B0–B3 step and compares them with the resolved condition definitions. For B2–B3, the allowed mechanism switch is institutionalization_enabled. The shared wiring installs the candidate generator, feature extractor, decision layer, execution operator, registry, model, and member-information interfaces for each condition.

An action-reachability test checks the offered candidates for the delivery chain: edit, commit, open, review, CI, merge, and release. It checks the chain in both B0 and B3. When the candidate-pool capacity limit binds, delivery actions are retained first. The conformance tests inspect this pool-priority rule and the shared governance role topology alongside the condition switches.

D.6.2 Configuration, Horizon, and Information Matching

The four conditions for a workload and seed are launched from one plan that fixes workload, horizon, seed, and model. Run records retain dataset_manifest_hash, starter_repo_digest, hidden_suite_hash, evaluator_environment_hash, tool_surface_fingerprint, ablation_fingerprint, llm_runtime_fingerprint, seed, and action_selection_mode. These fields bind results to the resolved configuration and support the workload/version/eligibility joins used by paired aggregation.

D.6.3 Member-Information and Access Tests

Access tests cover evaluator-private assets, future repository history, private member workspaces, and live-network access. Artifact, search, perception, and snapshot checks are specified in Appendix C.5.4. The optional prompt-visibility interceptor records the intercepted calls, and the future-recall checker inspects generated identifiers. The execution-layer proxy and socket controls in Appendix C.3.3 complement these member-facing checks.

E Internal Protocol Transfer and Executable Binding

The internal-transfer experiment uses 30 Text and 30 Exec GPT-5.6 Terra target runs over ten workloads and three seeds. Fresh uses the corresponding 30 B2 main-study runs. Text and Exec share a source-derived frozen rule package and vary its executable binding.

E.1 Protocol Source and Bundle Freezing

E.1.1 Source-Run Eligibility

The reported transfer package is constructed from B3-formed protocols using recurring capability families observed across source runs as the selection frame rather than target performance. From these families we selected functionally distinct representative concerns and normalized them into four source concerns and six closed-set typed runtime guards. This selection did not use target-evaluation outcomes. Normalization was limited to making rule semantics mechanically executable in the target runtime: timed, multi-head, or otherwise unobservable clauses were removed or rewritten; triggers, required fields/steps, affected actions, violation conditions, reason codes, and machine bindings were made explicit. Compatibility checks verified trigger observability and action reachability. Table 21 shows the public functional mapping.

The final package uses bundle schema org_capability_bundle_v2 and binding schema org_protocol_bindings_v2. It was frozen before the reported target evaluation. Text and Exec use exactly the same frozen readable content; the treatment contrast is the presence or absence of executable bindings.

Table 21. Public functional mapping of four source-protocol concerns to six compiled runtime guards.
Transferred protocol-to-guard mapping
Source concernCompiled guardAffected actionsFunctional role
S1issue owner assignedopen_prcustomer triage / ownership
S1branch owned by actoropen_prownership map
S2unchanged failed CI not retriedrun_ci, ci_testevidence workflow
S2 + S3current CI attestedmerge_prworkflow integration
S3independent reviewreview_pr, approve_pr, formal_pr_review, merge_prreview gate
S4release gate coveredpublish_product_releaserelease governance

Notes. S1–S4 identify four source concerns; the CI-attestation guard combines S2 and S3.

E.1.2 Package Provenance and Freeze Record

The transfer record identifies the package version, content digest, selected source items, normalization steps, and freeze time. A source-checkpoint contribution is linked to the corresponding item. The final freeze fixes the six compiled guards and their readable summaries before target evaluation, and Text and Exec reference the same package identity.

Source-selection metadata, source IDs, and normalization records are retained with the frozen package, connecting each compiled guard to the source concern in Table 21.

E.1.3 Target-Workload Selection

Targets use the W01–W10 starters, public requirements, and evaluators. Text and Exec are matched through their target/workload/seed/configuration mapping, yielding 30 matched units. Fresh uses the corresponding 30-run GPT-5.6 Terra B2 stratum. Seed rates are averaged within workloads, and applicable workloads receive equal weight.

Curation includes a reverse-reachability check: required verification steps are mapped back to agent-reachable artifacts and public issues before task admission. This check validates that the target action surface supplies the work needed to satisfy the protocol’s verification obligations. Source-to-target relations and scoring identities remain linked to the package and the target record.

E.2 Fresh-Member Initialization

E.2.1 Fresh Roster

The target organization is freshly initialized from the canonical eight-role roster. Member-local state is reset: private memory, personal workspace contents, sandbox cache, message read/acknowledgement state, commitments, reflections, wishes, and accumulated member-local experience are not inherited from the source run. Functional roles are reinstantiated for the target task, while the transferred organizational package is applied according to treatment.

E.3 Organizational-State Arms

E.3.1 Exec: Executable Protocol Binding

Exec initializes fresh target members and supplies the frozen package as readable rules with six live executable bindings. The target runtime receives the triggers, scopes, affected actions, required steps and evidence, responsibility functions, exceptions, and retirement conditions. Registration and identifier remapping associate subsequent runtime events with the imported rule and package.

Readable content is identical to Text. During target execution, existing bindings produce use and consequence records while the package remains frozen under the treatment in Appendix E.3.4.

E.3.2 Text: Content-Matched Protocol Presentation

Text initializes fresh target members and presents the same six ordered rule summaries to model-mediated work, including the code editor, through HOW THIS ORGANIZATION WORKS. The prose preserves each rule’s trigger, scope, required steps and evidence, responsible functions, exceptions, and consequences. The representation uses capability_as_prose and the TEXT_ONLY_DOC_TYPE carrier.

Text retains SDL and member learning with zero compiled bindings for the transferred package. The readable summaries and their order are checked against the same frozen package as Exec; the verbatim content is reproduced in Appendix I.7.1.

E.3.3 Fresh Reference: Reused B2 Main-Study Runs

Fresh uses the 30 GPT-5.6 Terra B2 main-study runs with ten workloads and three seeds, initialized without an inherited protocol package.

E.3.4 Frozen Formation and the B2 Reference

The treatment freezes the mechanism-producing path for the entire target window: no new protocol proposals, adoption of additional protocols, or learned rule revisions. Ordinary code edits, tests, communication, local memory, and the B2 member-learning paths continue. Imported executable bindings and their use/consequence records remain active while formation is frozen.

Table 22. Why the main-study GPT-5.6 Terra B2 condition is the fresh-start reference.
Transfer treatment freeze matrix
Target componentFresh referenceTextExec
B2 backbone: SDL and member learningsamesamesame
New protocol proposal / adoption / revisionoffoffoff
Fixed transferred protocol contentabsentpresentpresent
Transferred executable bindingsabsentabsentpresent
Origin of counted runsmain studytarget runtarget run

E.4 Transfer Payload

E.4.1 Permitted Payload

Only organizational state may cross: executable rules with their trigger_condition, scope, affected_actions, required_steps, required_fields, enforcement_rule, enforcement_action, exception_rule and sunset_rule; their responsibility mappings over functions; and the carrier documents those rules reference. The bundle is schema-versioned and validated on both export and import.

E.4.2 Excluded Payload

Excluded by construction: source code and patches from the source workload; any target patch or solution; private member memory; hidden-test outcomes and any evaluator record; task-specific answers; and future repository history.

E.4.3 Payload and Reset Records

Each transfer record contains a capability_transfer block with roster_origin, capability_form, source_repository_id, and protocols_injected. experiment_phase is capability_transfer, and pack identifies the source-to-target relation. Receipt validation checks the required fields, package contents, and source/target relation against the selected evaluation plan.

The record stores treatment identity, package-item and carrier-document counts, package digest, and reset state. A source-to-target role map binds obligations to target functions. Fresh initialization reconstructs member memory, workspaces, sandboxes, messages, commitments, reflections, and wishes while retaining the designated external protocol package. Runtime events refer to the injected package and rule identifiers, providing the join from inherited organizational state to later target actions.

E.5 Content Matching and Binding Control

E.5.1 Content-Matched Prose Construction

Text and Exec use the same canonical_v2 package. The six ordered summaries render to an identical 3,148-character code-editor block in both treatments. Exec activates six compiled bindings; Text activates none. Appendix I.7.1 provides the shared readable text and package identifiers.

E.5.2 Execution-Hook Removal

Exec registers the transferred package’s affected_actions matchers, responsible_roles routing hooks, and execution checks. Text retains the corresponding prose in shared work state with zero compiled bindings. Both treatments keep the target-stage proposal, adoption, and revision path disabled.

Conformance checks compare package identities, readable fields, and rendering order across Text and Exec and inspect the presence of compiled bindings in Exec and their absence in Text. Existing-binding use, violation, and execution events are recorded separately from protocol creation and revision, preserving the frozen target treatment throughout the observation window.

E.6 Transfer Estimands and Evaluation Scope

E.6.1 Primary Controlled Comparisons

Table 23. Internal-transfer comparisons.
Internal transfer: controlled contrasts
ContrastReferencePass-rate\Interpretation
Exec–TextActual target treatments+6.5 ppExecutable binding of content-matched protocols.
Exec–FreshReused GPT-5.6 Terra B2 equivalent+15.8 ppComparison with the designated no-package reference.

E.6.2 Interpretation

Exec–Text estimates the gain from executable binding of the content-matched source-derived package. Exec–Fresh measures the gain over fresh B2 initialization.

F Metrics and Statistical Analysis

F.1 Canonical Data Sources

F.1.1 Repository State and Action Log

Product and delivery outcomes use the repository objects in world.repo_system.repo, the product artifacts, and evaluator outputs for the corresponding candidate state. The action log records explicit action occurrences. Protocol lifecycle measures use the protocol ledger and proposal histories. These source interfaces retain the object identifiers needed to join a reported count to its underlying records.

The selected case labels repository and action counters separately and specifies the runtime paths connecting them in Table 48.

F.1.2 Evaluator Outputs

The final-evaluation artifact uses schema orgenv_oss_final_evaluation_v1. Each contract or check records task_id, command, owner, kind, base_status, candidate_status, required, baseline_repo_digest, candidate_repo_digest, execution_policy_hash, exit_code, stdout_hash, and stderr_hash.

Passing cases satisfy candidate_status == passed; a confirmed fix additionally satisfies base_status == failed. Checkpoint evaluation uses the same scoring fields at each saved step across the current 360-run study. Infrastructure and malformed-evaluation errors retain explicit statuses and enter the metric-eligibility rules in Appendix F.3.3.

F.1.3 Organizational-State and Protocol Records

Protocol measures join proposal-manager objects and their status histories with the protocol event ledger. Adoption, use, violation, enforcement, amendment, and retirement each have their own event type. Per-spec scalar counters provide a cross-check against the corresponding ledger totals.

Distinct affected targets are reconstructed from typed governed-object references in enforcement events. The references identify PRs, patches, commits, and other work objects, and link the object-level consequences to the protocol lineage that generated them.

F.1.4 Token and Compute Receipts

Provider receipts accumulate per call in usage_totals and are written to the run-level llm_usage record, including retry calls. Input, output, and cached-token counters are retained where supplied by the provider. Checkpoint token totals are stored with the corresponding saved state for the matched-spend analysis.

Run duration is recorded in seconds. Provider-token accounting and elapsed time are separate measurements; tool and container compute follows the resource-accounting definition in Appendix B.5.

F.2 Metric Dictionary

Table 24. Metric dictionary.
Metric dictionary
MetricNumeratorDenominatorBetterApplicability
Verified production (evaluator, mainline)
Evaluator-confirmed seeded issuesissues whose contract failed on starter and passes on candidateseeded issues in packhighall packs
Exposed cases on mainlineexposed units passingexposed unitshighpacks separating exposed from held-out
Held-out cases on mainlineheld-out units passingheld-out unitshighpacks with a held-out set
Complete contracts on mainlinehidden contracts fully passinghidden contractshighall packs
Workspace and delivery
Workspace contractscontracts passing on the workspace treehidden contractshighpacks with a workspace series
Workspace exposed casesexposed units passing on workspace treeexposed unitshighidem
Patch acceptanceaccepted patchesgenerated patcheshighall runs
Accepted work reaching mainlineaccepted patches present in a merged commitaccepted patcheshighall runs with patches
PR merge ratemerged pull requestsopened pull requestshighruns that opened a PR
Releasesrelease objects published— (count)n/aall runs
Organizational mechanism
Currently adopted protocolsregistry lineages with terminal adoption_status=adopted— (count)n/aall runs
Ever-adopted protocol lineagesdistinct run–protocol lineages with an adoption record, including later obsolete rules— (count)n/aruns with lifecycle records
Terminal protocol-registry objectsall registry objects, with adopted, proposed, and obsolete states counted separately— (count)n/aruns with a terminal registry
Mechanism usesrecorded use events, evidence level labelled— (count)n/aall runs
Runtime protocol enforcementsexported protocol-linked enforcement events; scope stated with the count— (count)n/aruns with the reported counter
Distinct affected targetsdistinct governed objects in enforcement events— (count)n/abundles with references
Efficiency
Tokens per tasktotal tokens summed over taskstaskslowall runs
Tokens per confirmed fixblock-level token totalblock-level confirmed-fix countlowpositive-fix model–workload blocks; macro averaged
Actions per outcometotal actionsconfirmed fixeslowidem
Diagnostic
Seeded completion (declared)tasks the board marks completeseeded issues—all packs
Completion credibilitydeclared issues also evaluator-confirmeddeclared-complete issueshighcommon eligible issue set; positive declarations

F.2.1 Verified Production Metrics

The four mainline production rates in Table 24 are computed from frozen evaluator records. Confirmed seeded issues pair a baseline-failing contract with a passing candidate status. Exposed and held-out rates use their own requirement sets and scoring units. A complete contract requires all constituent behavioral cases to pass, with partial credit retained at case level.

Each numerator is computed on the evaluated mainline artifact and paired with the corresponding frozen denominator. This procedure ties the reported production endpoints to both the starter behavior and the final integrated code.

F.2.2 Workspace and Delivery Metrics

Workspace metrics score local candidate trees; mainline metrics score the integrated shared repository. Delivery accounting follows generated patches, accepted patches, accepted patches contained in merged commits, opened pull requests, and merged pull requests. Repository objects supply these stages, while explicit action occurrences remain separate log counters.

Workspace/mainline comparisons align run identity, checkpoint or endpoint, and scoring units. Latency measures use recorded event or checkpoint times. The selected case additionally joins PR identities to their opening, verification, blocking, and merge records.

F.2.3 Organizational-Mechanism Metrics

Protocol measures include proposal, adoption, registration, activation, action matching, reading, use, enforcement, amendment, and retirement. Each event retains its protocol identity and time. The generic use event credits a successful action matching the protocol’s affected_actions; decision and execution effects have their own linked records.

The census identifies autonomous, seeded, and imported rules by origin. Revisions are grouped into the parent within-run lineage. Persistence is measured from a named lifecycle origin, such as adoption or activation, to the last qualifying use or recorded active step. Formation applies the unit, step, use-count, and 48-step criterion in Appendix A.4.3; the strong evidence grade joins third-party execution, governed-object changes, and evaluator outcomes.

F.2.4 Efficiency Metrics

Mean tokens per run averages provider-token totals over contributing runs. For cost per confirmed issue, tokens and confirmed fixes are pooled across the contributing seeds within each model–workload block. The block cost is its token total divided by its confirmed-fix total. Zero-success runs retain their expenditure within the block, and blocks with positive confirmed-fix counts supply the eligible block ratios.

Arm estimates macro-average their eligible block ratios: B0–B3 use 18, 21, 21, and 22 blocks, respectively. The paired B3–B2 contrast computes both conditions on the same 20 common blocks before aggregation. Table 38 retains the Terra+Opus auxiliary calculation on its 14 common blocks. Internal-transfer costs use the separate transfer analysis associated with Table 2.

F.2.5 Diagnostic Metrics

Completion credibility is the fraction of declared-complete issues confirmed by the evaluator.

F.3 Aggregation, Coverage, and Missingness

F.3.1 Metric-Specific Pooling

For each run rr and metric kk, the analysis records numerator nrkn_{rk}, denominator drkd_{rk}, applicability, and evaluation status. It first computes the eligible run-level rate nrk/drkn_{rk}/d_{rk}, averages the three seeds within each model–workload–condition block, and then equally weights applicable model–workload blocks across the three-model study. The Terra+Opus auxiliary analyses use the same block-macro procedure on their two-model population.

The numerator and denominator travel together with the metric’s scoring unit and contributing-run set. Measured zero outcomes retain their zero value; inapplicable or undefined rates retain their NA status. Paired contrasts apply the same metric-eligibility and workload/version mapping to both conditions before recomputing the aggregate.

F.3.2 Checkpoint and Endpoint Coverage

The reported checkpoint curves and main endpoints cover the 360-run three-model study. Curves use the 24-step grid, with tree type and metric-specific scoring units fixed for each series.

F.3.3 Missing and Inapplicable Metrics

The analysis retains metric applicability, denominator validity, evaluator status, run completion, and version eligibility as separate fields. na denotes an inapplicable or undefined rate, including held-out metrics for workloads without a held-out set and outcome-normalized ratios with a zero denominator. A measured failure contributes zero to the corresponding pass numerator while retaining its valid denominator.

Infrastructure and malformed-evaluation errors retain explicit error statuses. The recorded eligibility rule determines which metric rows enter the paired calculation, and both conditions use the same rule. Run and version records preserve the reason associated with an excluded or superseded observation.

F.3.4 Exact Count Records

Each reported metric is associated with its raw numerator, denominator, scoring unit, contributing-run count, and analysis-population identifiers. The aggregation records retain the workload and model weights used to produce the displayed rates. These fields provide the numerical inputs for recomputing the table entries and matching an aggregate back to its constituent run and evaluator records.

F.4 Estimands

F.4.1 Primary Organization Contrast

The primary contrast is B3–B2 within matched model–workload–seed units for verified production and resource outcomes. The design contains 30 matched units per model, or 90 across the three models.

F.4.2 Secondary Organization Contrasts

B2 minus B1 estimates the SDL-based configuration bundle, including selection mode, profile conditioning, and member learning (Appendix D.1.3); B1 minus B0 estimates having long-horizon multi-agent collaboration under the documented condition definitions. All three models cover the same ten-workload, three-seed, four-condition design. Both are reported as secondary ladder diagnostics.

F.4.3 Transfer Estimands

Exec–Text uses 30 matched GPT-5.6 Terra workload–seed units. Exec–Fresh uses the corresponding B2 main-study runs. Seed rates are averaged within workloads, then applicable workloads receive equal weight. Held-out outcomes use W02–W10, or 27 observations per treatment.

F.4.4 Heterogeneity Estimands

Heterogeneity is summarized by workload, workload family, seed, and model. Model-stratified endpoints cover all three models. Workload/family and leave-one-workload-out summaries use Terra+Opus, with run-level rates macro-averaged over the corresponding blocks.

F.5 Statistical Analysis and Supplementary Checks

F.5.1 Matched Effects and Clustering

The paired main-study unit is model–workload–seed. Bootstrap resampling holds the model–workload blocks fixed and samples seeds within each block. For a treatment contrast, corresponding seeds are resampled as pairs, and each draw recomputes the same block-macro estimate used by the endpoint table. Arm intervals recompute their respective arm estimates.

This procedure preserves the shared starter, requirement set, and evaluator associated with each workload while retaining the paired treatment structure. Population-specific resampling settings are given below.

F.5.2 Confidence Intervals and Bootstrap

The production endpoints in Table 1 use 50,000 bootstrap draws. Each arm interval recomputes that arm’s block-macro estimate; the B3–B2 interval resamples matched seeds within each fixed model–workload block.

The three-model outcome-normalized-cost analysis uses 20,000 draws. Arm intervals use the respective eligible blocks, and the cost contrast resamples matched seeds within the 20 common B2–B3 blocks. Terra+Opus workload/family, leave-one-workload-out, and budget-capped analyses use 10,000 draws with random seed 1729.

Internal-transfer intervals come from the transfer analysis associated with Table 2. The in-situ complete-contract interval uses 200,000 matched-seed draws within its ten workloads; Table 43 reports paired 95% intervals for all four in-situ endpoints.

F.5.3 Multiplicity

The complete-contract B3–B2 contrast is primary; secondary endpoints are reported with pointwise 95% intervals.

G Full Results and Robustness Checks

G.1 Run Inventory

G.1.1 Main-Study Coverage

The main study contains 120 runs per model and 30 runs per condition within each model. Table 20 summarizes the experiment inventory.

G.1.2 Transfer Runs

Table 25 summarizes the transfer treatments and reused Fresh reference.

Table 25. Internal-transfer treatment accounting.
Internal-transfer treatment accounting
ConditionRun accountingState and origin
Text30 new runsTen workloads × three seeds; fresh members and readable protocol content.
Exec30 new runsThe same ten-workload, three-seed design; fresh members and content-matched executable package.
Fresh reference30 reused runsGPT-5.6 Terra B2 main-study reference.

G.1.3 Case-Diagnostic Runs

The mechanism case uses W01, GPT-5.6 Terra, seed 1401, and the 336-step horizon. Appendix H gives the protocol and delivery timeline.

G.2 Workload, Seed, and Family Outcomes

G.2.1 Per-Workload Outcomes

Table 26 reports Terra+Opus workload-level B3–B2 effects using the block-macro aggregation in Appendix F.3.1.

Table 26. Per-workload mainline B3–B2 point estimates.
Per-workload verified production: B3–B2
WorkloadComplete\Confirmed\ issuesExposed\
W01-6.67 pp-6.67 pp-6.19 pp
W02+5.56 pp+4.76 pp+4.76 pp
W03+14.81 pp+19.05 pp+16.67 pp
W04+16.67 pp+16.67 pp+16.67 pp
W050.00 pp0.00 pp0.00 pp
W06+6.25 pp+8.33 pp+8.33 pp
W07+13.54 pp+18.06 pp+18.06 pp
W08+7.69 pp+9.26 pp+9.26 pp
W09+1.04 pp+1.39 pp+1.39 pp
W10+6.25 pp+8.33 pp+8.33 pp

G.2.2 Repair and Construction Split

Table 27. Construction and repository-based workloads reported separately.
Verified production by workload family
MetricB0B1B2B3B3–B2
Construction
Complete contracts on mainline6.92%8.10%9.25%15.32%+6.07 pp
Evaluator-confirmed seeded issues8.37%7.70%10.30%17.06%+6.76 pp
Repair
Complete contracts on mainline13.69%28.32%23.37%30.32%+6.95 pp
Evaluator-confirmed seeded issues18.33%38.24%31.39%40.46%+9.07 pp

Notes. Repository-based workloads W06–W10 include bug fixes and feature development.

G.3 Full Endpoint and Time-Series Results

G.3.1 Checkpoint Curves

Figure 4 pools checkpoint trajectories and endpoints across the three-model, 360-run study. Series are evaluated on the 24-step grid and rendered as step-held curves between checkpoints.

G.3.2 Workspace and Mainline Outcomes

Figure 4 presents workspace and mainline outcomes. Appendix H.3 traces repository delivery through seven 48-step windows in the mechanism case.

G.3.3 Declared Completion and Evaluator Confirmation

Figure 4a compares declared completion and evaluator confirmation across all three models. Dashed and solid curves use the same checkpoint grid and eligible issue set.

For one eligible run/checkpoint let DD be the set of seeded issues declared complete, EE the set confirmed by the evaluator, and II the common eligible issue set. The two marginal rates are ∣D∣/∣I∣|D|/|I| and ∣E∣/∣I∣|E|/|I|. The object-level summary reports declaration confirmation ∣D∩E∣/∣D∣|D\cap E|/|D|, unconfirmed declarations ∣D∖E∣/∣D∣|D\setminus E|/|D|, and confirmed-but-undeclared issues ∣E∖D∣|E\setminus D|.

Table 28. Object-matched declaration diagnostics.
Matched declaration diagnostics
Matched diagnosticB0B1B2B3
Declared issues confirmed, |D E|/|D|56.68%56.41%54.01%48.37%
Declared issues unconfirmed, |D E|/|D|43.32%43.59%45.99%51.63%

Notes. Terra+Opus endpoint declaration records.

G.4 Action Allocation and Delivery Conversion

G.4.1 Action Totals and Category Coverage

Table 29 reports mapped and unmapped action totals for the Terra+Opus runs. Action shares use mapped actions within each condition.

Table 29. Terra+Opus action-log totals.
Terra+Opus action-log accounting
Action accountingB0B1B2B3
Mapped actions9,45986,28882,11188,289
Unmapped actions0100

G.4.2 Delivery Funnel

Table 30. Raw delivery-object counts across 60 runs per condition.
Delivery funnel: repository-object counts
Stage (60-run raw totals)B0B1B2B3
Generated patches1,8689,0488,6108,079
Accepted patches1,5086,6846,4736,381
Accepted patches reaching mainline5101,4351,9712,234
Opened PRs4981,6952,3612,908
Merged PRs3521,1681,7432,090

Notes. Delivery rates use model–workload macro averaging.

G.4.3 No-Visible-Progress and Rework Diagnostics

No-visible-progress actions record effort or status without touching an artifact, patch, branch, review, or message. The action taxonomy assigns work_on_task, rest, defer, overtime, and idle to this category.

G.5 Organizational-Mechanism Inventory

G.5.1 Proposals and Adoption

Table 31 summarizes 617 terminal protocol objects across the 60 Terra+Opus B3 runs: 513 adopted, 92 proposed, and 12 obsolete. Adopted objects comprise 393 autonomous and 120 environment-seeded rules. The 12 obsolete lineages each retain a nonempty adoption_tick, which joins their earlier adoption to the terminal lifecycle state. Runtime protocol_enforcements counts accumulate over the recorded trajectory, including events preceding retirement.

Table 31. B3 protocol-state inventory and recorded runtime activity.
B3 protocol inventory and runtime activity
A. Terminal registry status across the Terra+Opus 60 B3 runs
Terminal statusAutonomousEnvironment-\
Currently adopted393120
Proposed, not adopted920
Obsolete after adoption120
Terminal objects by origin497120
All terminal protocol-registry objects617
All currently adopted protocols513
All ever-adopted lineages, including later obsolete rules525
B. Recorded runtime activity by workload
WorkloadRecorded usesRuntime\
W016,4504,199
W024,4531,270
W034,8484,297
W044,648582
W054,2313,997
W064,1552,452
W074,2354,373
W086,0875,144
W095,2301,970
W105,2535,159
All B3 runs (all protocols)49,59033,443
B0/B1/B2, each condition00

Notes. The environment-seeded rules are proto_review_before_merge and proto_experiment_logging, each present in all 60 Terra+Opus B3 runs. Runtime activity totals include both autonomous and seeded rules.

G.5.2 Use and Enforcement

Table 31 reports protocol-use and enforcement events. The mechanism case links 211 enforcements to 12 pull requests (Appendix H.8.1).

G.5.3 Amendment and Lifecycle

Of the 280 formed lineages, 182 have revisions, totaling 1,231 revision entries and 1,231 folded or superseded proposal identifiers. Retirement is recorded as obsolete; repair proposals use policy_repair_proposal.

G.5.4 Capability Families

Recurring protocol themes include commit-bound verification, interface-to-contract checking, review assignment, experiment logging, release readiness, and handoff readiness.

G.5.5 Corpus, Counting Units, and Formation Coverage

The autonomous-protocol census covers the 60 Terra+Opus B3 runs. Autonomous lineages exclude proto_experiment_logging and proto_review_before_merge.

The census counts within-run protocol lineages identified by (run, protocol_id), with revisions grouped into their parent lineage. Formation uses the criterion in Appendix A.4.3; the strong grade adds object-state and evaluator-linked evidence.

Coverage.

The census links terminal protocol states, structured rule content, formation records, family labels, roles, and revisions across all 60 Terra+Opus B3 runs. Each formed registry ID resolves to its rich ProtocolSpec; the (run, protocol_id) key joins its proposal, adoption, use, and revision records while grouping versions into one lineage.

Table 32. Autonomous-protocol census over the Terra+Opus B3 population.
Autonomous formation: Terra+Opus 60-run census
Stage or coverageAll 60\ 3 runsInterpretation
Autonomous proposal lineages497Excludes both seeded rules.
Currently adopted lineages393Terminal adoption_status=adopted.
Unadopted proposal lineages92Terminal adoption_status=proposed.
Subsequently obsolete lineages12Terminal obsolete; all have an adoption step.
Ever-adopted lineages405393 currently adopted plus 12 subsequently obsolete.
Weak-or-strong formed280Exact formation count.
weak-only248Strong is excluded from this row.
strong subset32Object-level and evaluator-linked evidence grade.

Across the complete 60-run census, 393 of 497 autonomous proposal lineages remain adopted at the endpoint (79.1%). Another 12 were adopted and later became obsolete, giving 405 ever-adopted lineages (81.5% of proposals); the remaining 92 are unadopted proposals. The 280 weak-or-strong formed lineages constitute 69.1% of the ever-adopted set and 56.3% of all autonomous proposal lineages. Of the formed lineages, 248 are weak-only (88.6%) and 32 strong (11.4%). Every formed registry ID maps to its rich ProtocolSpec, allowing versions to remain in one lineage.

G.5.6 Semantic Clustering and Functional Categories

We distinguish raw naming from functional grouping. The 280 formed lineages use 243 distinct raw protocol_type strings. These strings capture free naming and task-specific wording. All 280 map to a rich spec with an existing, runtime-recorded family. The present grouping uses these recorded labels, with content inspection organized around triggers, concrete obligations, roles, affected objects, and intended effects. Eight of the ten registered families occur; docs and other do not occur as primary family labels.

Table 33. Recorded functional families of formed autonomous protocols.
Formed capabilities by functional family
Recorded familyFormed\ShareRuns\/60Tasks\/10GPT-5.6\Claude\ 4.6Strong
Review and merge10838.6%4810387014
Release engineering8028.6%461039418
Evidence governance5620.0%381026306
Debugging176.1%1293144
Customer103.6%740100
Ownership72.5%66250
Experiment10.4%11010
Budget governance10.4%11010
Total280100%—1010817232

Notes. Family counts use one primary label per lineage. Run and workload coverage may overlap across families.

Review and merge.

The largest family typically responds to an opened or updated PR, or a claim that work is merge-ready or release-ready. Obligations include mapping issue acceptance criteria to tests, rerunning applicable CI against current mainline, recording a non-author review decision, and refusing readiness claims or merges when evidence, ownership, CI, or review is incomplete. The 108 lineages cover all ten workloads.

Release engineering.

These 80 lineages center on release candidates and release gates. A recurring sequence assigns an owner, locates the failing module, applies a patch, runs targeted and smoke/readiness tests, obtains reviewer sign-off, and reruns the release gate before publishing or declaring readiness. This family and review/merge both use CI and evidence, but differ in their principal object: a release candidate versus a PR and mainline promotion.

Evidence governance.

The 56 lineages regulate claims about fixes, compatibility, resolution, experiment results, or release status. They require applicable acceptance behavior, preserved compatibility, current verification status, and traceable evidence; insufficiently verified work remains pending, blocked, or under review rather than resolved or shipped. Coverage across all ten workloads shows recurrence of this protocol theme.

Debugging and recovery.

Seventeen lineages join failure signatures, suspected modules, patches, targeted tests, smoke results, and review sign-off into a closure chain. Four have strong evidence, a larger within-family fraction than the overall strong fraction.

Customer, ownership, experiment, and budget.

The ten customer lineages, all in the Claude Opus 4.6 records, connect customer issues to implementation or PRs, triage, owners, and response dates. Seven ownership lineages specify responsibility and handoff for tasks or artifacts. One experiment lineage requires shared tracker entries with run content, reproducibility state, seed/configuration, raw trace, and cost. The single budget-governance lineage targets institutional cost: when a protocol blocks too many requests in a week, its benefit and time cost must be reviewed and revision or retirement may be proposed.

Structural completeness, revisions, and role templates.

All 280 formed lineages have nonempty triggers, responsibility maps, success metrics, and enforcement-action fields. Problem evidence is present in 264 (94.3%), required steps in 274 (97.9%), and required fields in 270 (96.4%). These fields describe the structured content of formed protocols.

Table 34. Field completeness of the 280 formed rich protocol specifications.
Structured-field coverage of formed protocols
Structured fieldNonempty /280Coverage
Trigger condition280100%
Responsible roles280100%
Success metric280100%
Enforcement action280100%
Problem evidence26494.3%
Required steps27497.9%
Required fields27096.4%

Executor-role assignments are founder 179, cofounder 27, reliability 18, editorial 16, fast engineer 13, artifact design 12, community 8, and external voice 7. Every spec’s reviewer and approver fields contain the same cofounder + founder pair. The generation template fixes this reviewer/approver pair; executor-role assignments vary by protocol.

G.5.7 Diversity, Concentration, and Cross-Workload Reuse

The three largest families account for 244/280 formed lineages (87.1%). With family shares pkp_k, the observed Herfindahl index is ∑kpk2=0.276\sum_k p_k^2=0.276. Shannon entropy is H=−∑kpklog⁡pk=1.469H=-\sum_k p_k\log p_k=1.469 nats, giving H/log⁡8=0.706H/\log 8=0.706 across the eight observed families and an effective family count exp⁡(H)=4.34\exp(H)=4.34. The distribution has a long tail but is concentrated around review, release, and evidence obligations. These statistics summarize lineage-weighted diversity over the recorded functional families.

Table 35. Model strata of the Terra+Opus B3 protocol census.
Model strata of the capability census
Complete model stratumGPT-5.6\Claude\ 4.6
B3 runs with content and evidence3030
Formed lineages108172
Strong lineages1418
Mean formed per run3.605.73

The GPT-5.6 Terra stratum contains 108 formed lineages across five recorded families, while the Claude Opus 4.6 stratum contains 172 across eight. These strata characterize how the observed capability census is distributed across the two model settings.

Review/merge, release engineering, and evidence governance each occur across 10/10 workloads; debugging occurs across nine, ownership six, customer four, and experiment and budget one each. Within a recurring family, triggers, fields, and affected artifacts remain workload-specific.

Episode-window reconstruction shows post-adoption use beyond the originating episode window for all 280 formed lineages. This complements the cross-workload family recurrence and the independent-organizational-unit criterion used in the formation classifier.

G.5.8 Capability Coverage

The formed-protocol census is concentrated in review/merge, release engineering, and evidence governance, with smaller families covering debugging, customer handling, ownership, experiment tracking, and institutional cost. A complementary formation-funnel view follows visible friction through recognition, proposal, adoption, and sustained use, while ordinary repairs remain separate from newly formed organizational mechanisms.

G.6 Model-Conditioned Friction and Institutional Responses

We examine recorded work episodes and episode-to-protocol links under GPT-5.6 Terra, Claude Opus 4.6, and DeepSeek-V4-Flash, abbreviated as Terra, Opus 4.6, and DeepSeek.

Overlapping recorded pressures.

All three strata encounter debugging, launch pressure, experimentation, feedback, and claim disputes (Table 36). Typical debugging and launch-pressure counts are close, although other categories have different spreads. One Opus 4.6 PDF-workload run records 4,009 experiment episodes; retaining it raises that stratum’s experiment mean to 209.86, while its median is 26.5. We therefore describe typical activity using medians and interquartile ranges.

Table 36. Recorded friction and pressure episodes. Entries are the reported median [first quartile, third quartile] per run.
Episode categoryTerraOpus 4.6DeepSeek
Debugging33.0[27.5, 45.8]31.5[26.0, 39.8]29.5[24.0, 39.3]
Launch pressure34.0[33.0, 36.8]35.0[33.0, 37.8]33.0[32.0, 35.0]
Experiment activity27.0[24.3, 39.0]26.5[22.3, 33.0]21.0[18.0, 25.0]
Feedback ingestion12.0[10.3, 73.0]13.0[11.0, 37.5]13.0[11.0, 14.8]
Claim disputes1.0[0.0, 1.0]1.0[1.0, 1.0]1.0[0.0, 1.0]

Episode-to-protocol links.

Formed protocol specifications link source-episode categories to protocol families through source_episode_ids. Table 37 reports distinct linked lineages for selected source–family combinations. For example, Terra contributes 24 launch-pressure-to-release-engineering links; Opus 4.6 contributes 17 debugging-to-release-engineering links and 13 launch-pressure-to-review/merge links. DeepSeek contributes seven debugging-to-release-engineering links and five experiment-to-review/merge links. These records connect concrete work episodes to protocol formation.

Table 37. Recorded links from source episodes to formed protocols. Counts pool the three analysis strata.
Source episodeProtocol familyDistinct linked\
Launch pressureRelease engineering39
DebuggingRelease engineering34
Launch pressureReview / merge25
ExperimentReview / merge18
Feedback ingestionReview / merge14
DebuggingReview / merge11
Feedback ingestionEvidence governance5
ExperimentEvidence governance4
Launch pressureEvidence governance4
Claim disputeEvidence governance3
DebuggingDebugging3

Notes. A lineage may link to multiple source-episode categories.

Friction volume and institutional response.

Workload–seed-block-centered correlations vary by episode category. Debugging with debugging/release institutions has r=0.20r=0.20 (95% CI [−-0.07, 0.46]); claim disputes with evidence governance, r=0.19r=0.19 [−-0.15, 0.47]; and launch pressure with review/release, r=0.20r=0.20 [−-0.06, 0.46]. Experiment activity with experiment/evidence institutions is positively associated (r=0.24r=0.24 [0.08, 0.44]), while feedback with customer/ownership institutions is negatively associated (r=−0.24r=-0.24 [−-0.33, −-0.14]).

The episode links connect recorded work experiences to the protocol-formation history.

G.7 Cost and Efficiency Results

G.7.1 240-Run Token Accounting

Table 38. Terra+Opus token accounting over 240 runs.
Main-study token expenditure and outcome-normalized cost
Token metricB0B1B2B3B3–B2\
Terminal total (M)211.3092,634.198225.058391.372–
Mean / run (M)3.522[3.350,3.695]43.903[40.745,47.262]3.751[3.439,4.054]6.523[5.119,8.379]+2.772[1.373,4.665]
Per confirmed issue (M)4.947[3.396,5.963]31.760[25.994,46.592]2.907[2.402,4.632]3.651[2.359,6.660]-0.174[-1.686,0.633]

Notes. Brackets give 95% intervals. Mean-token/run uses all 60 runs per condition. Per-confirmed-issue arm estimates average each arm’s eligible nonzero model–workload blocks; the paired B3–B2 contrast is recomputed on 14 common nonzero blocks (42 paired runs), so it need not equal the arithmetic difference of the two arm estimates.

G.7.2 Outcome-Normalized Cost

Table 38 reports Terra+Opus cost estimates. Table 1 reports the pooled three-model estimates. Both use the block-macro cost definition in Appendix F.2.4.

G.7.3 Budget-Capped Checkpoint Sensitivity

For each (model,workload,seed)(\text{model},\text{workload},\text{seed}) pair, the budget-capped sensitivity takes B2’s final token expenditure as the cap and selects the last saved B3 checkpoint at or below that cap. The score is the evaluator output at the selected checkpoint. All 60 Terra+Opus matched units have a qualifying checkpoint with an evaluator output.

Token receipts provide recoverable monotone checkpoint curves for the 60 B3 runs, and terminal totals match the corresponding frozen token records. The analysis joins each checkpoint’s recorded cumulative spend, candidate state, and evaluator output. It equally weights model–workload blocks and uses 10,000 paired-seed bootstrap draws with seed 1729.

Table 39. Budget-capped checkpoint sensitivity.
Budget-capped checkpoint sensitivity
EndpointPairsBlocksB2B3B3–B2\[95% CI]
Complete contracts602016.310%23.760%+7.450 pp[3.491, 10.927]
Held-out cases541813.43%18.51%+5.08 pp[0.926, 11.111]
Exposed cases602022.570%34.290%+11.720 pp[6.781, 19.617]
Evaluator-confirmed seeded issues602020.840%27.052%+6.212 pp[4.528, 11.719]

Notes. Held-out outcomes use W02–W10; all other rows use W01–W10.

All four verified production contrasts remain positive at the selected budget-capped B3 checkpoints.

G.8 Complete Transfer Results

G.8.1 All Intervention Arms

Table 40. Complete internal-transfer summary.
Internal protocol transfer: complete outcomes and costs
MetricFreshTextExec
Behavioral-case pass rate25.4%34.6%41.2%
Exposed-case pass rate23.434%39.7%44.4%
Held-out-case pass rate11.111%11.1%25.9%
Complete-contract rate18.574%31.5%32.3%
Evaluator-confirmed issue rate21.720%41.4%45.1%
Pull requests opened, mean / run47.152.148.8
Pull requests merged, mean / run38.242.239.7
Releases, mean / run24.126.726.2
Patches generated / accepted, mean / run141.8 / 110.0124.7 / 122.4118.0 / 116.7
File-edit actions, mean / run141.9121.0118.0
Total actions, mean / run1,389.31,377.81,385.9
Imported mechanisms, mean / run06.0 (text)6.0 (executable)
New target mechanisms, mean / run000
Recorded uses, mean / run00588.0
Recorded binding enforcements, mean / run00197.0
Amendments, mean / run000
Mean tokens / run (M)2.3812.6152.448
Tokens / evaluated case (k)176154144
Tokens / verified pass (k)691444350

Notes. Count-valued rows are means per run. Confidence intervals for behavioral-case pass rate and cost metrics appear in Table 2.

G.8.2 Primary Controlled Contrasts

Exec improves behavioral-case pass rate by 6.5 pp [0.7, 15.6] over Text and 15.8 pp [9.2, 31.8] over Fresh. Exec also exceeds Text on each production endpoint in Table 40.

G.9 Robustness Checks

G.9.1 Equal-Task and Matched Subsets

Terra+Opus auxiliary analyses use ten workloads, three seeds, and B0–B3. Complete, exposed, and confirmed-issue contrasts have 60 pairs; held-out contrasts have 54 pairs over W02–W10.

G.9.2 Model-Stratified B3–B2 Effects

Table 41. Model-stratified B3–B2 effects.
Model-stratified B3–B2 effects
EndpointB2B3B3–B2\[95% CI]Applicable\
GPT-5.6 Terra
Complete contracts18.574%23.440%+4.866 pp[-1.713, 11.359]30/30
Held-out cases11.111%12.037%+0.926 pp[-4.630, 7.407]27/30
Exposed cases23.434%28.082%+4.648 pp[-0.011, 12.505]30/30
Confirmed seeded issues21.720%28.082%+6.362 pp[1.296, 14.220]30/30
Claude Opus 4.6
Complete contracts14.046%22.200%+8.154 pp[4.509, 13.271]30/30
Held-out cases15.741%35.185%+19.444 pp[0.167, 49.333]27/30
Exposed cases21.706%32.518%+10.812 pp[3.911, 17.613]30/30
Confirmed seeded issues19.960%29.438%+9.478 pp[5.113, 18.018]30/30
DeepSeek-V4-Flash
Complete contracts9.54%13.64%+4.10 pp[1.31, 7.13]30/30
Held-out cases1.85%4.63%+2.78 pp[0.93, 5.56]27/30
Exposed cases12.79%17.08%+4.29 pp[0.77, 7.92]30/30
Confirmed seeded issues12.05%17.08%+5.03 pp[1.73, 8.45]30/30

Notes. Each model has 30 matched pairs for complete, exposed, and confirmed-issue outcomes and 27 for held-out outcomes. DeepSeek intervals use 50,000 same-workload, matched-seed bootstrap draws.

All four B3–B2 point estimates are positive for GPT-5.6 Terra, Claude Opus 4.6, and DeepSeek-V4-Flash, with no endpoint direction reversal across models. For DeepSeek-V4-Flash, the paired 95% intervals for all four verified endpoints exclude zero. For held-out cases specifically, B0/B1 are 0.00% [0.00, 0.00], B2 is 1.85% [0.00, 3.70], and B3 is 4.63% [0.00, 8.33], with B3–B2 +2.78 pp [0.93, 5.56]. Its additional behavioral diagnostics are positive on mainline (B0/B1/B2/B3: 4.80 [2.78, 7.06] / 11.50 [9.42, 13.59] / 10.28 [7.09, 13.49] / 13.64 [11.55, 15.90]%; B3–B2 +3.35 pp [0.28, 6.59]) and on workspace (9.88 [6.98, 12.46] / 17.54 [15.20, 20.11] / 15.21 [11.50, 18.63] / 17.57 [14.76, 20.55]%; +2.37 pp [−-2.13, 7.15]). Under the same block-macro outcome-normalized cost definition used for the main-study paired estimand, the DeepSeek B3–B2 effect is −-1.087M tokens per confirmed issue.

G.9.3 Leave-One-Workload-Out Analysis

Table 42. Leave-one-workload-out robustness.
Leave-one-workload-out robustness
EndpointFull B3–B2\[95% CI]LOO >0LOO CI\>0
Complete contracts+6.51 pp[2.812, 10.517]10/1010/10
Exposed cases+7.73 pp[2.183, 13.016]10/108/10
Confirmed seeded issues+7.92 pp[2.908, 14.409]10/1010/10
4@l@ minipage[t] tabularx @L C0.258 C0.258 C0.142 @ 4lSensitivity range across workload omissions\ Endpoint & Minimum LOO & Maximum LOO & Max. abs.\ \* Complete contracts & +5.39 pp\ omit W04+7.98 pp\ omit W011.46 pp
Exposed cases+6.58 pp\ omit W07+9.27 pp\ omit W011.55 pp
Confirmed seeded issues+6.68 pp\ omit W03+9.54 pp\ omit W011.62 pp
tabularx minipage

Notes. Each refit omits one workload from both models and paired conditions. “LOO CI >0>0” counts intervals wholly above zero.

The three reported endpoints remain positive under every W01–W10 omission, so the aggregate direction is not attributable to a single workload. Complete-contract effects range from +5.39 to +7.98 percentage points, exposed effects from +6.58 to +9.27 points, and confirmed-issue effects from +6.68 to +9.54 points.

G.10 In-Situ Prose-Only Institutionalization Ablation

G.10.1 Design and Intervention

The internal Text/Exec experiment fixes a transferred package and disables new target-stage formation. Here we test executable binding inside the original organizational setting while proposal generation, governance, approval, adoption, and revision remain enabled. The completed ablation uses Claude Opus 4.6, all ten workloads W01–W10, and three seeds (1401, 2711, and 4013), yielding 30 new B3-text runs matched to the corresponding 30 executable B3 runs from the main study by workload and seed.

B3-text retains the B3 roster, member-local learning, reflection, wishes, protocol proposals, governance, adoption, revision, the core SDL scoring parameterization, and ordinary task and repository validity checks. The intervention changes what adoption produces operationally: an adopted ProtocolSpec remains a readable shared organizational rule and is rendered through the text-only decision-context interface, but it does not install an executable protocol binding. Learned-rule effects on selection and routing, automatic protocol-use recording, and protocol-specific enforcement are disabled. Environment-level CI, review, and transaction-validity checks continue to apply.

Both conditions begin at the initial state and independently develop their work, proposals, and adopted rules throughout the run.

G.10.2 Three-Seed Outcome Comparison

The primary in-situ endpoint is complete contracts on mainline. Complete, exposed, and confirmed-issue outcomes use 30 matched workload–seed units; held-out outcomes use 27 matched units because W01 has no held-out set. For each metric, run-level rates are first averaged over the three seeds within each applicable workload, then equally weighted across applicable workloads: ten for complete, exposed, and confirmed-issue outcomes, and nine for held-out outcomes. Table 43 reports the aggregate comparison.

Table 43. Three-seed in-situ B3-text versus executable B3.
In-situ binding ablation: verified mainline outcomes
EndpointB3-textB3B3-B3-textPaired 95% CI
Complete contracts on mainline15.03%22.20%+7.18 pp[+3.54, +10.91] pp
Held-out cases on mainline12.96%35.19%+22.22 pp[+10.85, +29.71] pp
Exposed cases on mainline18.47%32.52%+14.04 pp[+8.99, +20.23] pp
Evaluator-confirmed seeded issues17.72%29.44%+11.71 pp[+6.27, +15.91] pp

Notes. Rates macro-average the three seeds within each workload, then equally weight applicable workloads. Intervals are paired 95% CIs; held-out outcomes use W02–W10.

Executable B3 exceeds B3-text by 7.18 percentage points on complete contracts (95% CI [+3.54, +10.91]). The same direction appears on held-out (+22.22 pp), exposed (+14.04 pp), and evaluator-confirmed seeded-issue (+11.71 pp) endpoints. All four paired 95% intervals are above zero.

G.10.3 Matched Complete-Contract Endpoints

Table 44 reports complete-contract counts for the 30 matched B3-text / B3 pairs.

Table 44. Matched complete-contract endpoints for B3-text / B3.
Matched complete-contract endpoints: B3-text / B3
WorkloadSeed 1401Seed 2711Seed 4013
W01 Mini Blobstore0/5 / 0/50/5 / 0/50/5 / 0/5
W02 Traffic Watch2/9 / 2/90/9 / 2/91/9 / 1/9
W03 TG Automation1/9 / 5/91/9 / 1/91/9 / 3/9
W04 PDF Reformatter0/4 / 1/40/4 / 2/40/4 / 1/4
W05 FastAPI Dashboard0/4 / 0/40/4 / 0/40/4 / 0/4
W06 Boltons10/16 / 11/1610/16 / 12/1612/16 / 10/16
W07 Celery1/16 / 3/161/16 / 1/162/16 / 2/16
W08 Soup Sieve1/13 / 1/130/13 / 1/131/13 / 1/13
W09 cattrs4/16 / 5/162/16 / 4/164/16 / 2/16
W10 Tenacity6/16 / 4/162/16 / 5/165/16 / 3/16
10-workload macro17.23% / 25.42%10.49% / 22.85%17.37% / 18.34%

Notes. Each cell gives B3-text / executable B3. The three-seed aggregate is 15.03% / 22.20%, yielding the +7.18 pp contrast in Table 43.

G.10.4 Seed-Stratified Endpoints

Table 45 reports seed-stratified endpoint rates.

Table 45. Seed-stratified endpoint macros for the in-situ ablation.
Seed-stratified endpoint macros
MetricArmSeed 1401Seed 2711Seed 4013
Complete contractsB3-text17.23%10.49%17.37%
Complete contractsB325.42%22.85%18.34%
Held-out casesB3-text13.89%2.78%22.22%
Exposed casesB3-text20.96%13.10%21.37%
Evaluator-confirmed seeded issuesB3-text19.21%13.10%20.87%

G.10.5 Binding Manipulation Check

Table 46 reports proposal, adoption, readable-rule, and executable-binding state across the 30 matched runs.

Table 46. Three-seed B3-text manipulation check.
B3-text manipulation check
CheckSeed 1401Seed 2711Seed 4013B3-text totalB3 total
Protocol proposals106115103324347
Adopted protocols708176227321
Endpoint-renderable adopted rules668075221249
Cells with readable-rules block10/1010/1010/1030/30—
Cells with executable binding enabled0/100/100/100/3030/30
Recorded protocol-use events000025,079
Protocol-specific enforcement events000018,594

Notes. Endpoint inventories and whole-run event totals across the matched 30 runs; seed columns each cover ten workloads.

H Case-Study Evidence: Commit-Bound Verification

This case traces protocol formation, enforcement, revision, and delivery in the W01 Mini Blobstore run with GPT-5.6 Terra and seed 1401. Displayed member, protocol, and event identifiers use consistent aliases.

H.1 Case Configuration

H.1.1 Matched Diagnostic Configuration

Four runs, one per arm, on Mini Blobstore (W01) at seed 1401 with a 336-step horizon and model gpt-5.6-terra with low reasoning effort. All four share pack, starter tree, hidden suite, evaluator, seed, horizon, model, tool surface and information; circadian rhythm is ablated in all four. The primary endpoint:

Table 47. Matched case diagnostic, endpoint on mainline.
Selected case: verified mainline endpoints
EndpointB0B1B2B3
Behavioural cases passing0/350/350/3528/35
Complete contracts0/50/50/52/5

H.2 Repository and Action Counts

H.2.1 Repository-State Counters

Table 48. Repository objects and action occurrences in the matched case.
Selected case: repository and action counters
SourceCounterB0B1B2B3
repo statecode patches accepted74168113118
repo statepull requests opened181680
repo statepull requests merged02476
action logedit_repo_file75169131112
action logcommit_patch1125093
action logopen_pr00320
action logmerge_pr02128
action logrun_ci33427142126
CI recordsruns recorded73651589339

Notes. Some runtime paths create or merge pull requests without emitting a separate open_pr or merge_pr action-log entry.

H.3 Work Without Delivery

H.3.1 Local Production

The baseline arms continued to perform editing and verification actions. Actions taken per 48-step window are flat to the horizon: at t ∈ [0,48)t\,{\in}\,[0,48) the four arms take 23, 206, 194 and 208 actions; at [96,144)[96,144) they take 22, 198, 184 and 231; at [192,240)[192,240) they take 18, 204, 183 and 200; at [288,336)[288,336) they take 17, 185, 174 and 234. B1 is as busy at t=330t=330 as at t=10t=10. B0 is a single member, so its twenty per window is comparable effort per head.

For local editing, B3 made 112 edit_repo_file actions and had 118 code patches accepted, against B1’s 169 and 168. B3 has the fewest recorded file-edit actions of the three multi-member arms. The selected B1 mainline endpoint remains zero despite its larger number of these local operations.

H.3.2 Shared Delivery

Where they differ is handover. B0 committed once. B1 merged two pull requests, both inside the first 48 steps, opened its last at t=96t=96, and spent the remaining 240 steps writing and checking without handing anything over: its final 96 steps contain 114 CI runs, 63 edits, 62 internal searches, 51 public-test runs, and no pull request opened or merged. B2 opened sixteen and merged four, all four inside the first 48 steps; its final 96 steps went to work_on_task (114) and update_task_status (77), which is task bookkeeping around work that never left the desk. B3’s final 96 steps are shaped differently: work_on_task 50, run_ci 36, edit_repo_file 27, commit_patch 26, approve_proposal 18.

H.3.3 Delivery Over Time

Table 49. Pull requests opened / merged per 48-step window, from repository state differenced per window.
Selected case: shared delivery over time
WindowB0B1B2B3
steps 0–481 / 06 / 212 / 413 / 9
steps 48–960 / 02 / 03 / 013 / 12
steps 96–1440 / 00 / 01 / 013 / 14
steps 144–1920 / 00 / 00 / 07 / 9
steps 192–2400 / 00 / 00 / 013 / 9
steps 240–2880 / 00 / 00 / 010 / 15
steps 288–3360 / 00 / 00 / 011 / 8
Total1 / 08 / 216 / 480 / 76

Early B3 adoptions occur at t=21t=21, 3939, 5858, 6060, and 6565. All recorded baseline merges occur in the first 48-step window; baseline PR openings extend into later windows, through t=96t=96–144 for B2. Adoption timing, PR creation, and merged delivery are shown as distinct series.

H.4 Capability Inventory and Governance

H.4.1 All Proposed Mechanisms

Table 50. Every mechanism this run proposed, adopted or failed to adopt.
Selected case: complete proposal and adoption inventory
ProposalProposedAdoptedUsesEnf.Amend.
P01step 12step 214192112
P02step 13step 395800
P03step 7step 581800
P04step 48step 6024704
P05step 48step 6524207
P06step 264step 2763900
P07step 44never———
P08 (1st)step 267never———
P09 (2nd)step 278step 291———
P10 (3rd)step 321never———

Notes. Event-ledger proposal and adoption records. A dash denotes an unavailable count.

H.4.2 Adopted and Rejected Mechanisms

The run proposes ten mechanisms, adopts seven, and leaves three unadopted. A merge-readiness handoff is adopted on its second attempt. P02 uses a generic review-before-merge template; P03 uses a generic experiment-logging template.

H.4.3 Uses, Enforcements, and Amendments

Counts are in Table 50. At the t=336t=336 endpoint, elapsed time since adoption is 336−21=315336-21=315 steps for the principal rule and 336−60=276336-60=276 steps for the current-commit integration-evidence rule. Lifecycle records separately track activation, revision, retirement, and last-active state. Amendments total 13 across three protocols, the last at t=317t=317, so revision continued to within twenty steps of the horizon. The case’s 211 reported enforcements are linked to P01 and 12 distinct pull requests (Appendix H.8.1). P02–P06 have positive reported use counts and zero reported enforcements.

H.5 Related Rules for Verification and Evidence Freshness

H.5.1 Five Convergent Rules

Five related mechanisms address evidence completeness, current-commit verification, and readiness as repository state changes. Four members proposed them (M02, M04, M01, M05). The following clauses retain the source requirements, with task-identifying names and object identifiers anonymized; the clauses are:

  • P01 (M02, proposed step 12, adopted step 21, 315 steps from adoption to the endpoint): a pull request affecting public callables or cross-module behaviour is merge-eligible only when its review record identifies every changed public callable signature, identifies every added or modified cross-module invocation, maps each identified item to the applicable written contract, and records successful execution of the applicable public smoke tests through the highest stacking step affected by the change; anything unmappable, or any required smoke failure, must be resolved or explicitly rejected before merge approval.

  • P04 (M04, proposed step 48, adopted step 60, 276 steps from adoption to the endpoint, revised four times): a pull request changing a stacking-step implementation or a shared-state boundary must not be merged unless its recorded CI result and relevant public smoke result both identify the exact current PR head commit as the tested commit and both pass; if mainline changes after either result is recorded, the PR must be synchronized and both rerun on the resulting current head; and a reviewer or merger must be able to compare the recorded tested commit with the current head and determine pass or fail without relying on an author’s statement.

  • P05 (M01, adopted step 65, revised seven times): no release, launch announcement or readiness claim may proceed until the checklist is complete, CI passes on current mainline, required reviews are complete, and the external claim is limited to the verified dependency chain.

  • P06 (adopted step 276): not merge-ready until the review record contains a complete trace for each changed completion entry point and a linked passing affected integration or public-contract smoke result for the PR head.

  • P09 (adopted step 291 on the second attempt): the owner records current mainline revision and passing CI immediately before merge; any mainline movement invalidates that.

H.5.2 Shared Principle

The following formulation summarizes their shared verification concern:

A verification result belongs to a commit, not to a pull request — and when mainline moves, the result expires.

The integration-evidence rule explicitly binds both CI and public smoke results to the current PR head and requires reruns after mainline movement. The readiness rule refers to current-mainline evidence, while the contract-boundary rule emphasizes checkable interface-to-contract mappings. The trace and handoff rules add related obligations.

The learned rules organize responsibility, review evidence, and readiness obligations around the shared commit-bound CI workflow.

H.5.3 Relation to the Observed Failure Mode

The CI ledger measures the absence of that discipline in the other arms (Table 51): B0 made one commit and ran CI against it 73 times, with none passing; B1 made twelve commits, recorded 651 CI runs, one single commit accounting for 162 of them, and passed two. The baseline arms repeatedly tested a small set of commits while delivering few changes. This repeated-verification pattern coexists with the environment-level commit-binding checks shared across conditions and motivates the additional organizational evidence discipline learned in Relic.

H.6 CI Ledger

H.6.1 Runs, Commits, Passes, and Merges

Table 51. CI ledger for the matched diagnostic.
Selected case: commit-bound verification and delivery
ConditionCI runsCommitsRuns/\CI passesMerged
B073173.000
B16511254.222
B25895011.854
B3339933.612676

H.6.2 Repeated-Verification Pathology

B0 made one commit in 336 steps and ran CI against it 73 times; none passed. B1 made twelve commits and recorded 651 CI runs, of which its most-tested object, C-A, accounts for 162 by itself; two passed. B2 sits between, at 11.8 runs per commit and five passes. B3 tested each of 93 commits about three and a half times.

The distinction that matters is between testing a lot and testing something that has moved. A rule requiring current-head evidence and freshness after mainline movement expresses a discipline relevant to this pathology.

H.7 One Rule’s Governed Lifecycle

H.7.1 Originating Friction and Challenge

The initiating episode is EP-A, recorded as a speed-versus-quality tension in which the timeline did not establish that the implementation work had been validated end to end. The first challenge is a formal review in which M01 refuses pull request PR-C and states the requirement in the refusal: add the missing tests, provide the relevant test file and evidence, and request a review before it can proceed. Recurrence follows as repeated ci_contract_break events against different pull requests inside the same window. All three are events with identifiers, actors and steps, and all three were visible to the proposer at proposal time.

H.7.2 Proposal, Support, and Adoption

The proposal is made at t=12t=12, naming signature drift across a module boundary as the friction it is written against. It is supported and adopted at t=21t=21, with M02 as supporter. On adoption the compiled spec is registered and enters the action context.

H.7.3 First Use and First Enforcement

The first use event occurs at step 22, one step after adoption. At step 26, the ledger records a protocol-linked refusal of the requested merge and retains paired before/after snapshots of the governed PR. The event identity, step, protocol reference, and object snapshots connect the adopted rule to this concrete execution consequence. Subsequent PR records link the same object to repair, verification, and its final delivery state.

H.7.4 Amendments While in Force

The rule is amended twice while in force, at t=33t=33 by M01 and at t=59t=59 by M02. Governance applies to the amendments as well as to the rule: the source of the second, AM-A, was itself edited by M02 at t=39t=39 and by M01 at t=43t=43 before it was adopted. Enforcement per step rose from 0.50 to 0.68 across the first amendment and from 0.55 to 0.69 across the second. These rates summarize the rule’s recorded enforcement activity across the two amendment intervals.

H.7.5 Later Reuse

The principal rule records 419 uses and 211 enforcements through t=336t=336, spanning 315 steps from adoption to the endpoint. It satisfies the formation criterion in Appendix A.4.3. P09 is adopted at t=291t=291 and has a 45-step post-adoption observation window.

H.8 Pull-Request-Level Consequences

H.8.1 Distinct Pull Requests Affected

The 211 enforcement events involve 12 distinct pull requests.

H.8.2 Delayed and Later-Merged Work

Eleven of the twelve blocked requests subsequently merged, 15 to 112 steps after their first recorded block, usually with additional patches. These block-to-merge intervals include the subsequent implementation, verification, and coordination needed before delivery.

H.8.3 Work Still Blocked at the Observation Horizon

One request, PR-B, remained blocked and unmerged at the recorded horizon, without an observed override, despite approval by an agent occupying the reviewer role. Its head, C-B, was verified 87 times. The organization responded by building a tool named after the repair rather than by overriding the gate, preserving the rule while pursuing a repair.

H.8.4 Unaffected Merges

Of the 76 merges, 65 were never blocked by this rule. Eleven previously blocked requests later merged and one remained unmerged at the horizon.

H.9 Case Interpretation

The case traces recurring integration friction into a proposed and adopted rule, followed by repeated use, pull-request enforcement, amendments, and sustained shared delivery.

I Integrity and Reproducibility

This section connects the execution interfaces, frozen artifacts, run records, and recovery procedures used to reproduce and inspect the reported experiments.

I.1 Evaluator Isolation

I.1.1 Filesystem and Process Interfaces

The member-facing substrate exposes the project identifier, product name, dataset directory, and starter directory. Evaluator assets occupy the private hidden-test, held-out-requirement, and reference directories. Scoring creates a one-shot candidate workspace outside the simulation and hashes its tree before and after execution. The recorded hashes bind the score to the exported candidate state.

In the container configuration, evaluation runs in a separate untrusted, network-disabled container built from a digest-pinned image. In the host configuration, it runs in a separate process against the exported tree. The execution-policy receipt identifies the mode, dependencies, and artifact identities used for the corresponding score (Appendix C.3.3).

I.1.2 Hidden-Test and Reference Isolation

The evaluator resolves hidden suites and reference content through the 32-byte token-keyed vault described in Appendix C.3.2. World state and checkpoints retain the versioned binding. Constant-time token comparison controls vault lookup. Artifact, search, perception, and snapshot tests inspect the member-facing representations, and initialization validates that private evaluator assets use the vault interface.

I.1.3 Output Interface

Endpoint evaluation runs after rollout and writes evaluator-side contract/check records. During rollout, members receive the outputs of public tests, CI, and ordinary readiness/release gates, including their associated failure reasons. Hidden-test source, assertions, held-out requirements, and reference implementation remain in the private evaluator interface. The run records retain both the member-visible execution feedback and the offline evaluation needed to reconstruct the sequence.

I.2 Temporal and Network Controls

I.2.1 Network Restrictions

Repository operations use the in-process repository model, and the search system’s five domains use internal or frozen corpora. The frozen-web domain is a seeded document snapshot. Replay resolves searches through the saved corpus and cache; an unresolved lookup raises a cache-miss error. Product and evaluator execution use the proxy and loopback-only socket controls in Appendix C.3.3. Model-provider calls are the designated network egress. When enabled, the prompt-visibility interceptor records call hashes and coverage.

I.2.2 Future-History Restrictions

The member-visible starter is the designated frozen task snapshot. Public issue text supplies the problem and acceptance requirements. Later repository versions, hidden requirements, reference code, and source-identifying metadata are retained in evaluator-side provenance. Workload names and upstream version pairs in the paper link the results to those frozen source records.

I.2.3 Access Validation

Runtime access tests, enabled prompt audits, and post-run future-recall checks record information-interface validity. A flagged run is handled under the recorded exclusion policy. The flag, run identity, and validity reason remain associated with its execution and analysis record.

I.3 Benchmark Identifiers and Artifact Joins

W01–W10 consistently join workload manifests, result tables, transfer records, and case-study references. Table 16 gives their names, original pack IDs, and upstream version pairs in the frozen analysis order. Member, protocol, and event aliases retain consistent mappings within the case. These identifiers connect the paper’s summaries to the corresponding task definitions and event records.

I.4 Run-Bundle Schema

I.4.1 Configuration and Identity

The versioned record experiment_run_record.json uses schema orgenv_experiment_run_v2. Its identity fields include run_id, pack, condition, arm_id, seed, provider, model, and reasoning_effort. Configuration fields retain action_selection_mode, profile_assignment, experiment_phase, oss_control, evaluation_perturbation, and mechanism_ablations.

The record schema associates these values with the ablation_fingerprint, llm_runtime_identity, llm_runtime_fingerprint, resource_budget_fingerprint, tool_surface_fingerprint, and information_budget_fingerprint. Artifact fields include dataset_manifest_hash, starter_repo_digest, reference_repo_digest, candidate_repo_digest, hidden_suite_hash, evaluator_environment_hash, event_graph_hash, and run_manifest_hash. started_at, ended_at, and duration_seconds retain execution timing.

I.4.2 Execution Records

The bundle contains the action log, typed world-event log, model-call records, and perception-derived decision context. ActionDecision rows retain validation status and rejection reasons. Checkpoints retain member state and private memory, repository and product state, the protocol ledger and proposal objects, scheduler and background-job state, and the deep per-step replay frames and final snapshot.

Typed identifiers join actions to the work objects they changed and to the governance or protocol records relevant to the transition. The same joins support event-ledger counts, distinct-object counts, and case-study reconstruction.

I.4.3 Evaluation and Integrity Records

The final-evaluation artifact retains per-contract status and digest fields, and the reported checkpoint series covers the three-model study. llm_usage supplies provider-token receipts. organizational_capability_evidence retains formation evidence and its hash, while transfer runs carry a capability_transfer receipt. Enabled prompt audits use prompt_visibility_audit with a chained hash.

A completeness block checks the record against the expected field profile for its run and phase, and record_notes preserves validation details. Required-field counts, availability statuses, and artifact identities remain attached to the record. These checks provide an explicit input-validation step for replay and aggregation.

I.5 Checkpointing and Replay

I.5.1 Serialization

Whole-world serialization stores a versioned payload with a SHA-256, a separately hashed metadata sidecar, and the case-plan, source-provenance, model-binding, and resource-budget fingerprints. The live model client is detached before serialization and reattached after loading. The payload preserves the organization, member, scheduler, and background-job state needed to continue from the saved step.

I.5.2 Resume Equivalence

A zero-step resume loads the checkpoint and checks state equivalence before advancing the simulation. The resume path validates the case-plan, source-provenance, model-binding, and resource-budget fingerprints and rejects a mismatch. The supervised launcher uses this path for recovery after machine interruption.

I.5.3 Replay and Continued Execution

Recorded-event replay reconstructs the checkpoint world state, seeded environment randomness, saved search corpus and cache, and evaluator identity. Case-study timelines use the recorded events and checkpoints. Continued execution restores this state and then issues new provider calls under the saved configuration. The run records retain the restored checkpoint and subsequent calls as one execution lineage.

I.6 Run Validity and Recovery

I.6.1 Failure Taxonomy

Timeout, transport, empty/malformed response, and schema failures enter the provider-error handling path with bounded within-call retries. Returned-model counters preserve the observed model identity. Container and dependency faults are recorded as CI or public-test infrastructure errors. Qualification and evaluator failures retain contract-level statuses, and checkpoint scoring uses its per-contract timeout. Checkpoint hashes validate serialized artifacts; machine termination recovers from the last saved checkpoint.

I.6.2 Retry and Rerun Policy

Provider errors are retried within the call according to its deadline and retry policy. An interrupted run resumes from the last checkpoint under the same scenario–condition–seed identity, retaining earlier partial artifacts. A new statistical draw uses a new seed. A provider blackout activates the batch circuit breaker and pauses remaining launches. Superseded records retain their lineage and status while the designated completed record supplies the aggregate.

I.6.3 Run Inclusion and Version Mapping

The main endpoint matrix contains 360 condition runs. Each paired analysis applies its declared population, metric eligibility, and frozen asset/version mapping. Run identity, completion status, scoring eligibility, selected checkpoint lineage, and replacement reason are retained as separate analysis fields.

Starter or evaluator changes produce a new frozen asset identity. The aggregation record joins each retained outcome to its designated task, source package when applicable, manifest, evaluator, and completion status. Superseded or invalidated records remain linked through the provenance record, supporting reconstruction of the final included set.

I.7 Prompts, Schemas, and Configuration Artifacts

I.7.1 Exact Condition Prompts

Relic separates what-action selection from the model calls used to perform, reflect on, or formalize a selected activity. B0 and B1 use an LLM to choose from a constrained candidate menu. B2, B3, B3-text, Transfer Text, and Transfer Exec use SDL for what-action selection; their model calls are module-specific.

Shared cognitive-module guard.

For reflection, proposal drafting, proposal evaluation, and protocol synthesis, the system message combines the rendered workload/company brief, global grounding rules, the module name, and the module instruction. The common guard is:

You are the cognition of one agent in a simulated startup. You PROPOSE, REASON,
COMPOSE, EVALUATE and VERBALIZE. You do NOT change the world: the system
validates and executes. Never invent facts, objects, or actions that are not in
the provided context. Return ONLY valid JSON matching the schema.

The composition pattern is:

SYSTEM =
  COMPANY_BRIEF_TEMPLATE.format(company_name, product_name,
                                product_stage, product_purpose)
  + "\n\nGlobal rules:\n"
  + GLOBAL_GROUNDING_RULES
  + "\n\nCurrent module: " + module_name
  + "\n\n" + module_instruction

USER =
  agent_identity_for(agent, world)
  + "\n\n"
  + module_user(rendered_context)

B0/B1 action choice.

Choose the next intentional action for {agent_id} at tick {tick}.

{rendered_action_context}

Return ONLY JSON matching the action schema
(candidate_action and target_object_id MUST identify
one exact entry from candidate_options).

The model receives the available action menu, candidate options, valid targets, work state, recent events, task board, inbox, pull requests, test results, and other member-visible context. The returned action is validated against the candidate pool before execution.

Code editing.

After an edit action has been selected, the code editor receives the complete target file, public evidence available to the member, the edit goal, and readable organizational rules when present:

Target file:
{target_object_id} -- {title}
{summary}

CURRENT FILE CONTENT:
{complete_file_content}

{current_build_error, if present}
{published_contract, if present}
{imported_public_API, if present}
{readable_protocol_rules, if present}

Edit goal:
{edit_goal}

Reason:
{rationale}

Known gaps:
{known_gaps}

Return the patch JSON with `edits` = the list of anchored
replacements that make this change. Quote each `search`
exactly from the file above; do not return the whole file.

{revision_brief, if an earlier patch was refused}
{anchor_correction, if a replacement did not apply}

Reflection and wish creation.

The reflection user message is:

Reflect on the situation.
CONTEXT:
{JSON(reflection_context)}

Its full formation instruction is:

Reflect on the agent's own and the team's situation after recent events/episodes.
Output an honest self_assessment, team_assessment, blockers, and concrete
improvement_ideas (each with need_type/description/urgency/risk_if_unaddressed).
`need_type` says what kind of thing is missing, and the schema lists the kinds.
Choose protocol_need when the same avoidable mistake keeps happening and what is
missing is a standing rule everyone follows -- a checkable requirement on work
before it lands. Say in `description` what the rule should require and which
recurring failure it prevents; the organization writes and votes on the rule
from that. Choose policy_repair_need when an existing rule is the problem and
should be relaxed or repealed. The rules you are shown carry what each has been
holding up lately: a rule turning away most of the work while nothing ships is a
candidate however sound it reads, and one turning away nothing is not. Then say
in `description` how it should read instead -- write the replacement requirement
in full, keeping the part that was earning its keep and dropping the part
nothing can satisfy. A repair that only says the rule is too strict cannot be
applied: the wording you give is the wording work will be judged against
afterwards. When `what_has_been_failing` is present it is what the gates
actually said, in their words, with how the recent failures divide by kind.
Read it before you name a need: a failure that keeps arriving in the same words
is the clearest thing you have to reason from, and the need you name should
answer what those verdicts say, not what the situation feels like.

The main-study path maps and filters the structured improvement_ideas from this reflection output into wishes; there is no separate LLM wish-extraction call.

Wish →\rightarrow proposal.

The module instruction is:

Turn a wish into a concrete, feasible proposal composed from EXISTING
actions/capabilities/artifacts. Do not assume adoption -- name
approval_required_from. A wish is worded as a suggestion -- "propose a
checklist", "create a matrix". proposed_solution is not: it becomes the rule
the organization is held to, so write what must be true of a piece of work,
checkable by someone who was not there. Never write it as a plan to make a rule.

The user message is:

Draft a proposal for this wish.
CONTEXT:
{
  "wish": {visible wish record},
  "available_actions": [registered action identifiers],
  "product_context": {visible product context}
}

Recurrent-wish clustering.

The clustering system prompt is:

You read what the members of one organization have been asking for and name the
few themes that recur. A theme is worth a standing rule when the same avoidable
failure keeps costing the organization and a checkable requirement on work would
have caught it. Group the wishes: cite in `wish_numbers` the ones each theme
rests on, and name no theme that rests on fewer than two. Do not restate a rule
the organization has already adopted, and do not return two themes that a single
rule would cover -- each must stand on its own. `enforcement_rule` must be
checkable against a piece of work by someone who was not there: name what must be
true, not what people should care about. Return nothing rather than something
vague. Reply JSON only.

The user message provides already adopted rules and numbered wishes and asks for at most the configured number of themes worth a standing rule.

Pattern →\rightarrow candidate protocol.

Synthesize a candidate protocol from a repeated problem pattern. It is only a
proposal -- never adopt it yourself.

The corresponding user message provides the repeated pattern, occurrence count, candidate seed, and available actions.

Proposal evaluation.

Evaluate the proposal's feasibility/usefulness/risk/adoption. You only score and
suggest revisions; you cannot approve or reject.

with user message:

Evaluate this proposal.
PROPOSAL:
{JSON(proposal)}

The evaluator returns feasibility, usefulness, risk, adoption score, blocking issues, and suggested revision. Approval/adoption itself follows the governed review and approval path.

Wish-to-protocol formation criteria.

Table 52. Wish-to-protocol formation and measurement criteria.
StageRecorded criterion
Reflection → wishEach idea records need type, description, urgency, and risk; wishes retain source reflection, related objects/episodes, supporters, urgency, and status. Similar open wishes are merged.
Wish → proposalEligible when urgency is at least 0.85, it has at least two supporters, spans at least two episodes, or has founder endorsement; proposal provenance retains source wish/reflection/episode links.
Proposal → adoptionA protocol proposal with available provenance represents recurrence across at least two episodes or two wishes. Adoption requires at least three steps of review and two distinct approvers in the reported semi-automatic governance configuration.
Weak formationValid proposal/adoption, at least two post-adoption uses on distinct use units and steps, and at least 48 steps from adoption to the last qualifying use.
Strong evidenceWeak formation plus validated third-party enforcement, governed-state change, and independently evaluated outcome evidence.

Formation-record fields.

The per-run formation interface includes protocol_count, weak_protocol_count, strong_protocol_count, time_to_emergence, repeated_protocol_use_rate, cross_context_protocol_reuse_rate, protocol_persistence_rate, valid_third_party_enforcement_rate, measurable_impact, and protocol_proposal_rate. Post-adoption persistence, amendment/repair, and episode-transfer records provide lifecycle detail. The autonomous census joins these records by within-run lineage and applies the formation criteria in Table 52.

Readable protocol exposure in Transfer Text and Exec.

Both treatments use the frozen canonical_v2 package. Its semantic SHA-256 is a627adfcf6c299d0fabb185345bbc01106ea6901380aa0fbfee37c830ad8cdb6. The shared 3,148-character code-editor rule block has UTF-8 SHA-256 9c505e567afab8853436c1e7cbc609e51928de2e29e64be299d4721bef09ab49.

Table 53. Content matching for the Text/Exec transfer comparison.
Transfer componentTextExec
Frozen packagecanonical_v2canonical_v2
Six readable summaries and code-editor rendering orderidenticalidentical
Compiled machine bindings06
New target protocol formation/revisiondisableddisabled

The six verbatim readable summaries are:

  1. Issue ownership before PR opening. “Before open_pr, resolve the source branch exactly as the action handler does. If its commits link work, every explicit linked task must exist and have a non-empty owner_id. Every native Issue must have owner_id set, and every Task associated with that issue must also exist and be owned. A ProductArtifact issue has no inferred owner: it must resolve to at least one associated Task, and every associated Task must exist and be owned. A branch with no issue or task linkage is outside this guard. Assign every missing owner and retry; editing and committing remain allowed.”

  2. Branch ownership. “Before open_pr, resolve the source branch exactly as the action handler does and require actor.id to equal branch.owner_id. A non-owner is blocked without changing the branch; the recorded owner can open the request, so editing, committing, review, and merge repair lanes remain available.”

  3. CI retry on changed state. “Before run_ci or ci_test, resolve the pull request exactly as the action handler does. Block only when its latest CI result is failed and already attests the current source-branch head against the current mainline commit list. A new branch head or changed mainline permits a retry; a not_run infrastructure result follows the native retry path.”

  4. Current CI attestation at merge. “Before merge_pr for a pull request carrying commits or patches, require approved status and no merge conflict; require pr.commit_ids to exactly equal the source branch commit list and pr.patch_ids to exactly equal the ordered flattening of those commits’ patch_ids; require both pr.ci_passed and a latest passed CI result for the current branch head and require ci_base_main_commit_ids to equal the current repo.main_commit_ids. When the active profile exposes a deterministic candidate-tree digest, ci_tree_hash must exactly equal that digest; ordinary RepoLite uses the commit head and mainline base as its exact identity and does not require an unavailable digest. Missing, failed, conflicted, unsynchronized, or stale evidence blocks only merge; commit and run_ci remain available to repair it.”

  5. Independent review. “When the roster contains an agent other than the pull-request author, review_pr, approve_pr, and formal_pr_review by that author are blocked, and merge_pr requires at least one approved_by identifier that is both a current roster member and different from author_id. Unknown actor identifiers cannot review or approve. A one-agent roster may self-review so the workflow does not deadlock; request_changes and all repair actions remain allowed.”

  6. Release-gate coverage. “Before publish_product_release, resolve the release candidate exactly as the action handler does. Require approved status, the roster-aware release_approval_met check, no blockers, a non-empty included_pr_ids list whose requests all exist, are merged, and have ci_passed. Each candidate must carry the exact non-empty, duplicate-free gate profile selected by release_gates_for(world). Each profile gate must have exactly one recorded passed result or its exact identifier must already be listed in waived_gates; unknown gates, waivers, and results are rejected. Run readiness, collect the required approvals, and repair blockers before retrying publish.”

I.7.2 Configuration Artifacts

Frozen workload manifests and brief templates define the public task inputs, entry points, progressive contract steps, and public-test commands. Machine-readable schemas define the member and organizational state, proposal and protocol fields, and runtime bindings. The exact model prompts and transferred rule summaries above identify the text consumed by the corresponding cognitive modules.

Transfer metadata links the source package, content digest, freeze record, carrier documents, role mapping, and target reset state to the treatment configuration. Appendix E.4.3 specifies the capability_transfer fields used to inspect the injected package and its later execution records.

J Human–Organization Interfaces

Relic primarily studies whether autonomous agent organizations can accumulate persistent organizational capabilities. Once such an organization can maintain its own workstreams, responsibilities, review processes, protocols, evidence, and institutional state, however, a second problem appears: how should a human participate in an organization whose internal operating representation and execution tempo are increasingly machine-native?

We explore three interface designs for human participation in a continuously operating agent organization.

J.1 Why Direct Human Participation Becomes Difficult

Our initial design treated the human as another organizational member. This preserves an appealing symmetry: human and agent members communicate through the same channels, observe organizational objects available to their roles, and participate in the same workflows. In practice, this design exposes a representation mismatch.

Relic maintains structured organizational state including task ownership, review and CI status, active protocols, approvals, episodes, evidence, commitments, requests, and institutional memory. These representations are useful for agent decision making and organizational governance, but they are not necessarily the representation in which a human wants to understand ongoing work. A person entering the organization may first have to determine which workstreams matter, what is blocked, which evidence is current, who owns a decision, and whether an apparent problem has already been handled elsewhere.

Conversely, a natural-language message from a human is only one event in a continuously operating organization. Asking “Why is this blocked?” does not itself guarantee that the organization will stop, that the relevant agent will immediately answer, or that no other authorized member will act before the human finishes considering the issue.

A second prototype introduced a secretary agent that could help the human delegate implementation and verification work. This reduced some execution burden, but much of the underlying organizational state remained directly exposed. The human still had to interpret the machine-native organization and decide which internal objects deserved attention.

Figure 7 summarizes the interface progression that motivated a stronger human–organization abstraction boundary.

Interface progression before P3. In P1, the human joins the autonomous organization directly and must interpret machine-native organizational state. In P2, a secretary assists with delegated work, but substantial organizational complexity remains directly exposed. These experiences motivate a liaison-first abstraction boundary in P3.

J.2 Moving the Human–Organization Abstraction Boundary

We organize the design space into three interface levels. The progression changes how organizational complexity is presented to the human.

P1: Direct organization.

The human directly navigates organizational state and communicates with organizational members. This maximizes direct visibility into the organization but places the burden of orientation, interpretation, and attention allocation on the human.

P2: Transparent liaison.

A secretary assists with communication and delegated work while the underlying agents, workstreams, and organizational structures remain directly visible. The secretary reduces some execution burden, but the human still operates close to the machine-native representation and remains responsible for understanding a substantial fraction of organizational state.

P3: Liaison-first interaction.

The secretary becomes the primary human-facing interface. The organization is not simplified internally. Instead, its state is summarized, translated, and selectively surfaced to the human. Detailed organizational evidence and provenance remain available on demand.

Table 54. Three human–organization interface levels.
Hybrid human–organization interface levels
Primary interfaceMain interaction property
P1Organization directlyHuman interprets organizational state and communicates with members directly.
P2Secretary + visible organizationSecretary assists, but substantial organizational complexity remains exposed.
P3Secretary first; organization on demandSecretary translates state and intent while preserving access to evidence and provenance.

J.3 P3: A Liaison-First Hybrid Human–Agent Organization

P3 places the secretary at the human-facing boundary while preserving the autonomous organization. It combines three concurrently operating components: a full autonomous Relic organization, a human member with a human-facing office, and the shared organizational workflow.

Autonomous Relic organization.

The upper-right component of Figure 8 is the autonomous Relic organization. It continues to maintain its own workstreams, tasks, responsibilities, review processes, protocols, evidence, and institutional memory. Other organizational members continue working independently of the human-facing interaction.

The organization may discover problems, aggregate evidence, form recommendations, execute unrelated work, and update its institutional state without waiting for continuous human direction.

Human office.

The human is represented as a member of the larger organization and retains their own execution capacity. The human office contains the human, the secretary, and delegated execution agents that can carry out work belonging to the human member’s responsibilities. For example, a human may specify a desired behavior and constraints, after which the secretary can coordinate implementation or verification work on the human’s behalf.

The secretary consequently has two distinct responsibilities:

  1. organizational liaison: translate organizational state into human-readable summaries and route human intent back into organizational objects; and

  2. human-work coordinator: decompose and delegate work that belongs to the human member’s own organizational responsibilities.

The secretary coordinates the human-facing interface and delegated work while shared governance and the organization’s persistent decision processes continue to govern organization-wide work.

Shared organizational workflow.

Artifacts produced through the human office re-enter the same organizational workflow as work produced by other members. Code, documents, tests, and other artifacts remain subject to applicable review, CI, evidence requirements, and active protocols.

Together, these components form a hybrid human–agent organization: the agent organization continues operating autonomously, while the human participates through a human-facing abstraction layer and retains delegated execution capacity.

P3: a liaison-first hybrid human–agent organization. The autonomous Relic organization continues operating independently (upper right). The human interacts primarily through a secretary, which summarizes organizational state, surfaces items requiring human judgment, and translates human intent into organizational work. The human also has delegated execution agents inside a human office. Human-owned artifacts subsequently enter the same shared review, CI, protocol, and institutional workflow as other organizational work.

J.4 Bidirectional Translation

The secretary acts as a bidirectional abstraction layer between two different representations of work.

Organization to human.

The autonomous organization may contain many simultaneously active tasks, reviews, failures, protocols, and commitments. The secretary compresses this state into human-facing updates. We distinguish three forms of communication:

  • informational updates, for work that the organization has handled without requiring human action;

  • decision requests, for cases in which human judgment, preference, authorization, or unavailable information is required; and

  • on-demand explanation, through which the human can request the evidence, provenance, review history, or organizational trace behind a summary.

P3 therefore uses progressive disclosure. The default representation is concise, while the underlying organizational state remains inspectable when the human needs additional evidence.

Human to organization.

Human input is expressed in ordinary language but may imply different organizational actions. For example, “I do not trust this result” may indicate a request for additional evidence, whereas “always review these changes” may express a persistent policy-level intention rather than a one-time task.

The secretary can translate such input into structured organizational objects such as tasks, concerns, review requests, constraints, acceptance criteria, or protocol proposals. The resulting object then follows the organization’s ordinary validation, approval, and execution path.

J.5 Human Attention as an Organizational Resource

A highly autonomous organization should not require the human to inspect every internal action. At the same time, aggressive filtering can conceal important disagreement or make consequential decisions difficult to inspect.

This creates an organizational problem of deciding what deserves human attention. Routine and reversible issues may be handled internally; recurring but non-urgent issues may be summarized periodically; high-impact, irreversible, or unresolved decisions may require direct escalation. Escalation policy therefore becomes part of the human–organization interface rather than only a notification preference.

The completed walkthroughs expose a second dimension of the same problem: when human attention can arrive. An organization must decide not only which events deserve human judgment, but also whether its own execution should wait, slow down, continue autonomously, or defer irreversible actions while that judgment is pending.

J.6 Formative Paired Walkthrough Protocol

Four members of the author team completed formative paired walkthroughs of P2 and P3 after implementation. None had participated in the design or implementation of Relic or the two interfaces. Each explored ongoing work, inspected blockers and evidence, communicated a requested change, and responded to decisions requiring human judgment.

J.7 Paired Internal Walkthrough Results

P3 reduced organizational interpretation burden.

Across the four paired walkthroughs, the liaison-first P3 interface was consistently easier to follow than P2. In P2, the secretary could assist with delegated work, but the evaluator still had to inspect and interpret substantial machine-native organizational state before deciding how to intervene. Active workstreams, agent activity, pending reviews, protocol state, evidence, and other internal objects remained directly salient to the human interaction.

In P3, the secretary absorbed much of this interpretation burden. It summarized ongoing organizational state, surfaced the subset of issues that required attention, and allowed the evaluator to request underlying evidence or provenance when needed. The organization itself remained unchanged internally; what changed was the human-facing representation. The interaction therefore shifted from navigating the organization toward judging the decisions and exceptions that the organization surfaced.

Human deliberation became the remaining bottleneck.

A recurring observation was that improved legibility did not remove the difference in operating speed between humans and autonomous agents. While an evaluator was reading a summary, inspecting evidence, or considering whether to approve, reject, or redirect work, other agents could continue operating elsewhere in the organization.

In some interactions, useful work advanced while the human was still deliberating. When another agent had sufficient authority over a pending item, that item could also be resolved before the human responded. The resulting friction was therefore no longer primarily about understanding the state of the organization. It was about whether human judgment could arrive on the same timescale as the organization that was waiting for, or acting around, that judgment.

This makes the timing of intervention a separate interface variable. A human may understand the issue perfectly and still be too slow to participate in every decision at machine execution speed.

Pausing provides control but changes the workflow.

The interface can pause the organization while the human deliberates. This is a practical control mechanism: it prevents the relevant organizational state from continuing to evolve while the human reads evidence and forms a judgment.

However, pausing also changes the interaction regime. A continuously operating autonomous organization becomes temporarily synchronized to human decision speed. If every consequential decision requires such a pause, the resulting system behaves less like an autonomous organization with intermittent human oversight and more like a workflow that repeatedly waits for a human operator.

Representation mismatch and temporal mismatch are distinct.

The P2/P3 comparison separates two problems that can otherwise be conflated.

The first is representation mismatch: directly exposing machine-native organizational state makes it costly for a human to understand what is happening. P3 substantially reduces this burden through liaison-first summarization and progressive disclosure.

The second is temporal mismatch: even when the relevant state is presented clearly, human deliberation may remain slower than the organization that is acting around that decision. Better summarization reduces the time needed to understand a situation, but it does not make human reasoning operate at agent execution speed.

A human–organization interface must therefore decide both what should be escalated and how the organization should behave while judgment is pending. Possible policies include waiting only for irreversible actions, continuing reversible work, slowing selected workstreams, assigning decision deadlines, or allowing another authorized member to proceed after a timeout.

J.8 Authority and Interaction Records

Visibility, permissions, and decision rights follow the human and secretary roles. Human-owned artifacts re-enter the shared review, CI, and protocol workflow. Interaction records connect the human’s original instruction, the secretary’s structured interpretation, the resulting work or governance object, and subsequent execution.

The human can request the organizational evidence and provenance behind a summary. This record preserves the path from a human-facing decision to its source evidence and organizational consequences while the shared governance process remains active.

K External Software Benchmark: CooperBench

CooperBench evaluates joint implementation of interacting features in existing software repositories (Khatua et al. 2026). We evaluate the complete two-member Relic architecture on the full benchmark against released Solo, Peer/coop+git, and hierarchical Team trajectories. We additionally evaluate a fixed 48-pair same-model subset as a controlled test of coordination loss and recovery.

K.1 Full-Benchmark Scope and Success Criterion

The full CooperBench evaluation space contains 652 exact interacting-feature pairs. A pair succeeds when the delivered candidate passes the official tests for both features (both_passed=true); each exact pair is one scoring unit.

Benchmark-validity auditing ran alongside Relic execution. Once an issue family was established, all affected exact pairs were entered into the validity ledger. The completed audit excludes 183 exact pairs and leaves 469 valid pairs. Relic passes 371 of those 469 pairs; the Raw and definite-defect matched analyses use their corresponding exact-pair identities.

We report three validity policies:

  • Raw. No benchmark-validity exclusion is applied. The matched Relic analysis uses the 616 exact pairs with formal official verdicts.

  • Definite-defect exclusion. We remove only pairs affected by a directly established specification–verification mismatch, unpublished mandatory interface or representation, explicit behavioral conflict, or established cross-feature incompatibility. This policy excludes 111 unique exact pairs from 21 issue families, leaving 541 pairs.

  • Final validity exclusion. We remove every exact pair admitted by the completed benchmark-validity audit. The final ledger excludes 183 unique exact pairs, leaving a 469-pair validity-filtered benchmark.

For released systems, trajectories for all 652 pair identities are available. We therefore recompute their results on the same fixed 652-, 541-, and 469-pair universes. Only an official PASS contributes to the numerator; a non-passing or missing formal verdict does not further shrink the policy-defined denominator.

K.2 Two-Member Relic Adaptation

The adapter places two feature owners in one shared organization. Each starts from its assigned feature brief and works through a private workspace and branch. Requirements and local findings reach the other member through explicit messages or shared artifacts. Pair-specific run identities keep workspaces, patches, messages, and organizational state separate across concurrent pairs.

The B3-2 configuration retains SDL, governed protocol formation, and executable bindings. SDL selects feasible work, communication, review, and governance actions; the model produces code and messages. Members can propose rules from visible coordination failures, submit them for validation and approval, and apply adopted rules to subsequent work.

Joint delivery follows owner-local editing and public testing, commits and pull requests bound to a code revision, peer review, integration replay, and synchronization to shared mainline. Both members attest to the same mainline digest and export the same joint patch to their submission slots. The official evaluator then applies the two native feature test suites to that delivered candidate.

Public-probe receipts retain the probe, associated public requirements, candidate and relevant baseline statuses, and bounded execution errors and output. Fixed patch guards and artifact-identity checks implement the adapter, while learned protocols retain their own proposal, adoption, and execution records. The final official evaluation is linked to the exact pair identity, model, execution settings, and adapter revision.

Initial adapter development used a separate task to check coordination, integration, and submission before the reported evaluations. The full-benchmark Relic runs use Claude Opus 4.6 with high reasoning effort.

K.3 Baseline Configurations

The full-benchmark comparison uses three released GPT-5.5-hao references  (CooperBench 2026a, 2026b). Solo assigns both interacting features to one coding agent. Peer / coop+git assigns one feature to each of two peers without a permanent lead. Team no-proto uses the released lead–member hierarchy with shared task state and scratchpad.

Across the released runs, these structures exhibit a stable coordination hierarchy: Peer<Solo<Team.\text{Peer} < \text{Solo} < \text{Team}. Splitting interacting features across plain peers reduces success relative to Solo, whereas the structured lead–member Team exceeds Solo. Relic retains the peer topology with no permanent lead and adds persistent organizational state and executable governance to that structure.

The separate 48-pair controlled comparison uses Claude Opus 4.6 with high reasoning effort for Relic, Official Peer, and Official Solo. Official Peer assigns one interacting feature to each peer; Official Solo assigns both features to one agent. Relic retains the same model and two-peer feature split.

K.4 Benchmark-Validity Audit

K.4.1 Hidden Tests Versus Hidden Requirements

We distinguish withheld test instances from withheld requirements. A benchmark may hide concrete test code, input instances, edge cases, reference implementations, and evaluator data. However, for a hidden evaluation to measure satisfaction of the stated task, the acceptance semantics exercised by those tests must be supported by the public task contract available to the evaluated system. In short, test instances may be hidden, but task requirements should not be.

For example, if a public specification requires correct handling of all valid JSON inputs, a previously unseen valid JSON instance is an appropriate hidden test. By contrast, if the public specification requires JSON and CSV support while the evaluator additionally requires YAML, the evaluator has introduced an unpublished requirement rather than merely an unseen test instance. Likewise, a requirement to produce an error log does not by itself justify an evaluator that accepts only one unpublished exact log string.

Our audit therefore asks whether behavior enforced by the official evaluator is supported by the public specification.

K.4.2 Audit Scope and Outcome Distribution

The audit identifies 38 specification–verification issue families affecting 183 distinct exact feature pairs. The sum of family-level pair incidences is 194 because 11 exact pairs instantiate two issue families; deduplication by exact pair identity yields 183 affected scoring units.

Once an issue family is admitted, it applies to every affected exact pair. In the complete 183-pair final exclusion set, the raw Relic states comprise 21 PASS, 123 FAIL, and 39 pairs without an official verdict. The 111-pair definite-defect subset contains 5 PASS, 76 FAIL, and 30 pairs without an official verdict. The validity audit asks whether a scoring unit can distinguish failure to implement the public requirement from failure to satisfy an unpublished or materially under-specified evaluator assumption.

All original evaluator outcomes and experimental artifacts are retained. Validity exclusion changes only whether an exact pair contributes to the corresponding aggregate score; it does not rewrite its raw outcome.

K.4.3 Validity Evidence Levels

The six failure classes below describe what kind of contract problem is present. Independently, each issue family is assigned one of two evidence levels: a definite defect or a specification ambiguity.

Definite defects.

A definite defect is directly established from the public specification and evaluator behavior: an explicit specification–verification mismatch, an unpublished mandatory interface or internal structure, an unpublished exact representation, an explicit behavioral conflict, or an established cross-feature incompatibility. The audit contains 21 definite-defect families covering 111 unique exact pairs.

Specification ambiguities.

A specification ambiguity occurs when more than one behavior remains consistent with the public contract, while the evaluator accepts only one such interpretation. The D/A-tiered audit contains 12 ambiguity families covering 71 unique exact pairs; seven of these pairs also appear in the definite-defect set. The completed audit contains five additional issue families covering eight further exact pairs, for 38 issue families and 183 excluded exact pairs in total.

Thus, the definite-defect sensitivity policy excludes 21 families / 111 pairs, while the completed final validity policy excludes 38 families / 183 pairs.

Table 55. CooperBench validity taxonomy and evidence levels. ``Issue families'' counts all 38 root issue families in the completed audit. ``Primary pairs'' assigns each affected pair to one class in a fixed C1–C6 order so that the column sums to 183. D/A coverage reports the explicitly tiered definite-defect and ambiguity subsets and is not additive across classes.
ClassSpecification–verification failureIssue familiesPrimary pairsD coverageA coverage
C1Public contract and evaluator acceptance are inconsistent948470
C2Hidden interface, internal structure, or input-domain requirement734305
C3Callback protocol is not defined by the public contract318020
C4Behavioral scope or boundary semantics are not determined1149348
C5Exact output representation is not publicly specified320210
C6Cross-feature contract or retained-test incompatibility514120
Across-class total / union3818311171

The completed audit adds one C1 issue, one C2 issue, two C4 issues, and one C6 issue beyond the D/A-tiered family set; these five families account for the eight additional excluded exact pairs.

K.4.4 Complete Issue-Family Inventory

Table 56 lists every admitted issue family. Because CooperBench scores exact feature pairs rather than isolated features, a defect associated with one repeatedly paired feature can propagate to several scoring units.

Table 56. Complete CooperBench specification–verification issue-family audit. D denotes a definite defect; A denotes specification ambiguity; ``–'' marks a final exclusion family without a separate D/A sensitivity-tier assignment. ``Pairs'' is the number of exact feature pairs affected by the family and may overlap across rows.
IDClass / tierRepository / task / featureObserved contract problemPairs
R01C1 / Dgo_chi / 56 / f1The public example permits ``GET, HEAD'' as one Allow-header value, whereas the evaluator indexes two separate values from Header.Values("Allow").4
R02C2 / Dgo_chi / 27 / f3The evaluator directly requires internal logging symbols that are not part of the publicly specified interface.3
R03C4 / Ago_chi / 26 / f3The public contract does not fully determine the version-selection entry point or fallback behavior for an unknown selector or role; the evaluator selects one specific behavior.3
R04C4 / Aoutlines / 1371 / f2The public contract does not specify the behavior for unregistering a nonexistent filter; the evaluator requires unregister_filter("non_existent") to succeed without raising.3
R05C4 / Doutlines / 1371 / f4The public task specifies JSON/CSV behavior, while the evaluator additionally requires YAML support.3
R06C4 / Aoutlines / 1655 / f4The public contract permits colon or hyphen separators but does not determine whether they may be mixed; the evaluator rejects mixed forms.9
R07C6 / Doutlines / 1706 / f1The new feature requires wording that includes both Transformers and MLXLM, while retained tests from interacting features still match the older Transformers-only wording.7
R08C1 / Dtiktoken / 0 / f9The public contract names the parameter compress, whereas the evaluator calls compression.9
R09C2 / Dtiktoken / 0 / f10The public task specifies LRU/use_cache behavior, while the evaluator additionally requires an assignable enc.cache attribute that is not published in the contract.9
R10C3 / Atiktoken / 0 / f7The validator protocol is underspecified: the contract does not determine whether validation returns a Boolean or raises directly, while the evaluator fixes a False-returning predicate protocol.9
R11C5 / Dclick / 2068 / f10The public task requires error handling, but the evaluator additionally requires the unpublished exact text ``Editing failed''.11
R12C1 / Dclick / 2956 / f7The public example/model input uses an underscore-form parameter name, whereas the hidden acceptance condition requires the hyphenated CLI spelling.7
R13C4 / Aclick / 2956 / f3The public contract specifies ValueError for negative batch size but does not determine the exception class for non-multiple input; the evaluator requires TypeError.7
R14C6 / Dclick / 2068 / paired featuresOne feature explicitly requires wait(timeout=timeout), while retained mocks in interacting-feature tests expose a narrower wait signature.3
R15C6 / Dclick / 2068 / f1–f8The new feature moves to edit_files, while the interacting feature's retained tests still depend on the older internal edit_file API.1
R16C1 / Djinja / 1559 / f3The public task requires retries beginning at priority 0, while the evaluator also requires the priority-0 case to invoke the operation only once.9
R17C4 / Ajinja / 1465 / f5The public task discusses None, empty-string, and falsy values but does not define grouping semantics for a completely missing attribute; the evaluator requires one particular treatment.9
R18C3 / AHF datasets / 3997 / f2,f3The custom_criteria callback arguments are not specified. The evaluator accepts a single feature argument but rejects the alternative (name, feature) protocol.7
R19C4 / AHF datasets / 3997 / f4The public contract does not specify whether a JSON-formatted string supplied as an update value should be automatically deserialized into a dictionary; the evaluator requires that behavior.4
R20C3 / Apillow / 68 / f4The public task specifies an error_handler callable and high-level strategy but not its exact arguments or return protocol; the evaluator fixes (error,key,value) and specific fallback/skip semantics.4
R21C5 / Dpillow / 68 / f5The public task requires relevant logging, whereas the evaluator accepts a particular unpublished sentence form such as ``Orientation value: 3''.4
R22C1 / Dpillow / 290 / f4The public contract permits fallback to the maximum color count when the MSE threshold cannot be met, whereas the evaluator imposes strict error ordering. The reference image remains above the stated threshold even at 256 colors.4
R23C5 / Dllama_index / 17244 / f4The public force_mimetype contract does not require the exact data:image/png;base64, representation, while the evaluator does.6
R24C4 / Allama_index / 18813 / f5The public task mentions trimming but does not define it as removal of zero bytes from both ends of the raw binary payload; the evaluator fixes that interpretation.5
R25C2 / Dllama_index / 18813 / f6For format=mp3 without an explicit mimetype, the evaluator requires unpublished BytesIO.mimetype and BytesIO.as_base64 attributes.5
R26C2 / Ddirty_equals / 43 / f6The public task defines an issuer parameter, while test collection directly requires an unpublished class-subscript API such as IsCreditCard["Visa"].8
R27C4 / Adirty_equals / 43 / f7The public contract does not determine whether an invalid hash algorithm should fail at construction or evaluate to false during comparison; the evaluator requires a constructor-time ValueError.8
R28C2 / Dreact_hook_form / 153 / f6The public contract does not make the otherwise required onValid argument optional, while the evaluator supplies undefined and requires successful handling.5
R29C1 / Dreact_hook_form / 153 / f5The public task explicitly permits either a configuration object or an additional parameter, while the evaluator accepts only a third-argument options object.5
R30C1 / Dreact_hook_form / 85 / f2The public contract permits either formState.isLoading or a dedicated state representation, while the evaluator accepts only the former.4
R31C2 / Adspy / 8635 / f4The public task permits None to be intercepted at the call site, while the evaluator directly invokes a private helper with None; the private-helper input contract is not specified.5
R32C1 / Ddspy / 8587 / f3The public task describes timestamps as optional, while the evaluator requires specific private _t0/_t_last mappings and fixed output keys.5
R33C6 / Dpillow / 25 / f1–f2Feature 1 requires readonly state to be preserved when saving to a different file, whereas Feature 2 requires the default preserve_readonly=False behavior to ensure that the destination is writable before saving. The same default save-as operation therefore cannot simultaneously satisfy both public feature contracts.1
R34C2 / –click / 2068 / f5–f8The public feature permits lock-conflict failure but does not specify the exception class; the evaluator recognizes only a specific exception type.1
R35C4 / –pillow / 25 / f1–f5, f2–f5, f3–f5The public corner-marking contract does not specify source-image immutability or restoration after a marked save, while the evaluator enforces that boundary.3
R36C4 / –dirty_equals / 43 / f1–f3The public email contract does not specify a minimum final-label length, while the evaluator rejects the one-character final label used by this exact pair.1
R37C1 / –dspy / 8394 / f2–f3The public namespace contract requires the same value to be identical to None after assignment and remain callable as a context manager.1
R38C6 / –dspy / 8635 / f1–f6, f5–f6The paired public contracts impose incompatible adapter/probe revision and composition requirements across the interacting features.2

Why 38 issues affect 183 scoring units.

CooperBench scores exact feature pairs rather than isolated features. Consequently, a contract defect associated with one repeatedly paired feature propagates to every scoring unit containing that feature. For example, the unpublished exact-output requirement in click/2068/f10 affects 11 pairs; the tiktoken parameter-name and hidden-cache issues each affect nine; the jinja priority conflict affects nine; and the outlines separator ambiguity affects nine.

The 183 excluded pairs arise from 38 shared feature- or pair-level contract issues propagated through CooperBench’s pairwise composition.

K.5 Released-Baseline Sensitivity

The definite-defect policy leaves 652−111=541652-111=541 exact pairs, while the final validity policy leaves 652−183=469652-183=469. We apply the identical pair sets to all released trajectories.

Table 57. Released GPT-5.5-hao CooperBench trajectories under three benchmark-validity policies. Only official PASS outcomes enter the numerator; the denominator is fixed by the validity policy.
Validity policyNGPT-5.5 SoloPeer / coop+gitTeam no-proto
Raw652362/652 (55.5%)329/652 (50.5%)403/652 (61.8%)
Definite-defect541339/541 (62.7%)305/541 (56.4%)380/541 (70.2%)
Final validity469311/469 (66.3%)277/469 (59.1%)349/469 (74.4%)

The validity policy changes absolute pass rates but not the structural ordering: Peer / coop+git<Solo<Team no-proto.\text{Peer / coop+git} < \text{Solo} < \text{Team no-proto}. In the released CooperBench settings, plain peer cooperation is therefore the weakest of the three structures: splitting the interacting features across peers reduces performance relative to Solo, whereas the structured lead–member Team exceeds Solo.

K.6 Relic Results and Matched-Pair Sensitivity

Within each validity policy, Relic, Solo, Peer, and Team are compared on exactly the same exact-pair identities with formal Relic verdicts.

The final validity comparison contains all 469 valid exact-pair identities; the Raw and definite-defect sensitivity rows use their corresponding matched identity sets.

Table 58. Matched-pair CooperBench sensitivity. Within each row, all systems are scored on exactly the same pair identities with formal Relic verdicts under that policy.
Validity policyRelicPeerSoloTeam no-proto
Raw (n=616)388/616 (63.0%)317/616 (51.5%)350/616 (56.8%)391/616 (63.5%)
Definite-defect (n=535)383/535 (71.6%)303/535 (56.6%)337/535 (63.0%)378/535 (70.7%)
Final validity (n=469)371/469 (79.1%)277/469 (59.1%)311/469 (66.3%)349/469 (74.4%)

Raw comparison: lifting peer coordination to the Team regime.

On the unfiltered matched set, Relic reaches 63.0% on 616 identities, compared with 51.5% for the released Peer reference and 56.8% for Solo on the same identities. This is a +11.5 percentage-point gain over Peer and a +6.2-point gain over Solo, with Relic 0.5 percentage points below the hierarchical Team reference on the same pairs (63.5%).

The released CooperBench baseline ordering is Peer<Solo<Team.\text{Peer} < \text{Solo} < \text{Team}. Peer/coop+git is the weakest released coordination setting, whereas the lead–member Team is the strongest. Relic retains two peer feature owners with no permanent lead and reaches essentially the same performance regime as the specialized hierarchical Team on the unfiltered matched set.

Sensitivity to benchmark validity.

Under the definite-defect policy, Relic reaches 71.6%, compared with 56.6% for Peer, 63.0% for Solo, and 70.7% for Team on the same 535 pair identities. The corresponding differences are +15.0 points over Peer, +8.6 points over Solo, and +0.9 points over Team.

Under the final validity policy, Relic reaches 79.1% on the complete 469-pair validity set, compared with 59.1% for Peer, 66.3% for Solo, and 74.4% for Team. The corresponding differences are +20.0 points over Peer, +12.8 points over Solo, and +4.7 points over the hierarchical Team reference.

Relic’s advantage increases under stricter validity policies. On the raw matched set it exceeds Peer by 11.5 points and Solo by 6.2 points while remaining within 0.5 points of Team; under the definite-defect and final-validity policies it exceeds Team by 0.9 and 4.7 points, respectively.

From the weakest released topology to Team-level performance.

The released CooperBench hierarchy shows a coordination penalty for plain peers: splitting interacting features across two peers performs worse than assigning both to one agent, while a specialized lead–member hierarchy is required to exceed Solo.

Relic keeps the peer feature-owner structure and adds persistent organizational state, governed protocol formation, and executable coordination rules. This moves the peer topology from the weakest released CooperBench setting to Team-level performance on the unfiltered matched set, above Team under the definite-defect policy, and further above the hierarchical reference on the final validity set.

K.7 Frozen 48-Pair Same-Model Coordination Check

The same-model coordination check uses a fixed 48-pair subset spanning 30 task instances and 12 repositories, selected independently of Relic outcomes. The final audit excludes one of those identities, leaving 47 valid pairs scored for Relic, Official Peer, and Official Solo.

Table 59. Frozen CooperBench same-model evaluation pairs.
#RepositoryTaskPair
1dottxt_ai_outlines1371F1/F2
2dottxt_ai_outlines1655F1/F3
3dottxt_ai_outlines1655F6/F7
4dottxt_ai_outlines1655F7/F10
5dspy8563F1/F4
6go_chi27F3/F4
7go_chi27F2/F4
8huggingface_datasets7309F1/F2
9llama_index17070F1/F2
10llama_index17244F5/F6
11llama_index17244F2/F6
12openai_tiktoken0F4/F8
13openai_tiktoken0F1/F5
14pallets_click2800F1/F4
15pallets_click2800F1/F2
16pallets_jinja1621F6/F10
17pallets_jinja1621F1/F6
18pillow25F1/F5^
19pillow25F1/F4
20pillow68F1/F5
21pillow290F3/F5
22pillow290F2/F3
23react_hook_form153F2/F6
24react_hook_form153F1/F3
25samuelcolvin_dirty_equals43F2/F3
26samuelcolvin_dirty_equals43F2/F4
27typst6554F2/F6
28typst6554F1/F3
29dottxt_ai_outlines1706F4/F6
30dottxt_ai_outlines1706F5/F6
31dspy8394F3/F4
32dspy8394F3/F5
33dspy8587F1/F4
34dspy8587F2/F3
35dspy8635F1/F4
36dspy8635F4/F6
37go_chi26F1/F2
38go_chi26F2/F4
39go_chi56F1/F5
40go_chi56F2/F3
41huggingface_datasets3997F1/F2
42huggingface_datasets6252F4/F6
43llama_index18813F1/F5
44pallets_click2068F1/F4
45pallets_click2956F1/F8
46pallets_jinja1465F1/F7
47pallets_jinja1559F1/F8
48react_hook_form85F3/F4

†^{\dagger}pillow / 25 / F1/F5 was part of the original frozen 48-pair selection but is classified as BROKEN by the final benchmark-validity audit. We therefore exclude it from the same-model aggregate, exactly as broken pairs are excluded from the full-benchmark comparison. The reported controlled comparison consequently uses 47 final-valid pairs; this excluded identity was a Relic PASS and an Official Solo/Peer FAIL.

K.7.1 Same-Model Results

Table 60. Same-model CooperBench coordination check on 47 final-valid pairs from the frozen 48-pair subset.
SystemModelStructurePass / 47
Relic (B3-2)Claude Opus 4.6 (high)Two peer feature owners; no fixed lead28/47
Official SoloClaude Opus 4.6 (high)One agent receives both interacting features26/47
Official PeerClaude Opus 4.6 (high)Two peers split the interacting features13/47

Official Peer drops from Solo’s 26/47 to 13/47 when the interacting features are divided across peers. Relic retains the same model, reasoning setting, two-peer feature split, and absence of a permanent lead, yet reaches 28/47. Thus Relic not only recovers the coordination loss observed in Official Peer but exceeds the same-model Solo reference on this controlled subset.

Together with the full-benchmark comparison, the frozen-subset experiment provides a direct same-model control over the peer coordination setting, while the full benchmark evaluates the same peer-organizational design at benchmark scale against the released Solo, Peer, and hierarchical Team references.

L External Software Benchmark: ProgramBench

ProgramBench (Yang et al. 2026) evaluates a frozen executable protocol layer attached to the native Bash-action interface of a single mini-SWE-agent coding agent.

L.1 System Configuration

The baseline is official mini-SWE-agent (Yang et al. 2024). The treatment adds a stateful protocol sidecar around the same single-agent Bash interface. The model chooses actions, edits the candidate code, maintains the conversation history, and decides when to submit. Before dispatch, the sidecar checks the proposed action and accumulated execution evidence. Its response allows the action, requests current evidence, or refuses a transition that violates an executable rule.

The treatment uses mini-SWE-agent 2.4.6, ProgramBench 1.2.4, and GPT-5.6 with xhigh reasoning effort and all-turn context. Protocol content remains frozen throughout the target runs. Binding rules and advisory rules are identified separately in Table 62.

L.2 Task Selection and Fixed Evaluation Set

Adapter integration was developed on a separate demonstration task. We randomly sampled 25 ProgramBench tasks, fixed their identities, and evaluated both systems on this paired set.

Table 61. ProgramBench evaluation set. The task identifiers are randomly sampled once and then held fixed for both systems.
Fixed random sample of 25 ProgramBench tasks
1.\ zip-password-finder10.\ tig19.\ muffet
2.\ tui-journal11.\ parqeye20.\ elfcat
3.\ oranda12.\ igrep21.\ svd2rust
4.\ scc13.\ loop22.\ argc
5.\ rustowl14.\ dropbear23.\ code-minimap
6.\ datasurgeon15.\ bartib24.\ tty-clock
7.\ xh16.\ gdal25.\ nomino
8.\ rust-sloth17.\ proj
9.\ 7zip18.\ quinn

The sample contains 17 Rust tasks, three C++ tasks, three C tasks, and two Go tasks.

L.3 Frozen Protocol Package and Runtime Binding

The treatment uses the fixed programbench_pack_v0 package. It contains six executable binding rules and three advisory rules. Harness adaptation maps their triggers, evidence fields, and actions to the native mini-SWE-agent interface; the package remains fixed during evaluation.

Table 62. ProgramBench protocol package. Binding rules can affect execution; advisory rules are recorded as shadow guidance and do not themselves block an action.
RuleModeRuntime intent
PB0BindingRequire an original candidate implementation rather than copying, linking, or delegating execution to the reference executable.
P0BindingProtect public probes and other harness-scored surfaces from modification by the candidate.
P3BindingInvalidate stale verification after source changes; modified code must be rebuilt or rechecked in its current state.
P4BindingPreserve build/check failure semantics; an unsuccessful or unavailable check cannot be represented as successful evidence.
P7BindingRequire patch.txt, when present, to correspond to the current candidate code state.
P8BindingGate final submission on originality, non-empty source changes, current build evidence, absence of known failing current checks, and inspection of the current diff.
P1AdvisoryEncourage probing the available reference behavior on real inputs before repeated implementation changes.
P5AdvisoryEncourage obtaining new reference evidence when the same failure recurs without new information.
P6AdvisoryEncourage periodic re-inspection of the current source and diff during long trajectories.

L.4 Runtime Footprint

Across the 25 treatment trajectories, the coding agent executes 2,215 Bash actions. The protocol layer records 933 validation receipts and 815 inspections of source, diffs, or reference behavior, while issuing 40 hard refusals. Twenty-six hard refusals arise from P3, requiring current verification after the code state changes, and fourteen arise from P4, preserving a failed or unavailable build/check as a failure rather than accepting it as successful evidence.

The three advisory rules additionally record 293 shadow cases in which their recommended condition would have requested a different next step: 239 under P6, 37 under P1, and 17 under P5. These shadow events do not block execution.

The 933 validation receipts comprise 370 PASS, 48 FAIL, and 515 INCONCLUSIVE outcomes.

L.5 Scoring and Outcomes

We compute each task’s behavioral pass percentage using ProgramBench’s native evaluator, excluding tests marked ignored by that evaluator, and macro-average over the fixed 25 tasks. Both systems use the same scoring procedure.

Official mini-SWE-agent reaches 64.164%; the protocol treatment reaches 70.916%, an absolute gain of 6.752 percentage points (10.5% relative). All 25 treatment trajectories produce submissions and completed evaluations.