What Happened
AI Agents · Governance · Software Engineering · Benchmarking · Research Briefpaper of 20 Aug 2026 · package read at version 0.2.72, 28 Aug
arXiv:2608.20622 · a proposal, and the package it names
The Harness Ships, the Governance Doesn’t
On 20 August 2026 George Salapa posted a 21-page architecture proposal to arXiv: give a large enterprise one coding-agent harness, run it unmodified behind every surface, and the whole fleet becomes auditable, because reviewing a solution stops being a code review and becomes reading one instructions file. The paper names its reference implementation, which is a real Python package with 83 releases behind it. Reading that package is what this report does. The harness the paper describes is there in detail. The four services its governance argument rests on are not — and its licence forbids the forking on which the deployment model turns.
The figures Each one counted from a record anyone can fetch: the arXiv API, the PyPI API, and the package’s own source distribution
- 83releases behind the one URL the paper offers as its disclosure. It names none of them
- 0times
risky— the flag the paper’s whole risk mechanism turns on — appears anywhere in the shipped package - 8/8arXiv references that resolve, every author list matching the record in order — one of them 42 names long
- 0benchmarks in the paper, which says so itself: “No benchmark accompanies these claims”
The record, graded Sorted by how well each stands up, not by how interesting it is
| Standing | What | How it was checked |
|---|---|---|
| Confirmed | A 21-page architecture proposal was submitted to arXiv at 23:44 UTC on 20 August 2026, in cs.AI and cs.SE, by one author, with nine references and three figures | The arXiv API record and the abstract page agree with the document |
| Confirmed | All eight arXiv references exist, and every author list in the reference section matches the record exactly and in order — including a 42-author entry | Each identifier queried against the arXiv API and compared name by name |
| Confirmed | The reference implementation it names is real and current: 83 releases from 5 April to 28 August 2026, Python source included in the distribution | The PyPI JSON record, then the source distribution downloaded and listed |
| Confirmed | Its licence permits installing and running it for internal business use, and forbids copying, modifying, forking, decompiling, reverse engineering or creating derivative works without written permission | The LICENSE file inside the distribution; the PyPI classifier reads License :: Other/Proprietary License |
| Confirmed | The paper names no version. Sixty-nine releases preceded it and fourteen have followed, so the URL it cites has not meant the same thing twice | Upload timestamps for all 83 releases, split at the paper’s submission time |
| Confirmed | Nothing in the package names the governance layer: risky, gateway, entitlement, config.yaml and live_skills each appear zero times | Case-insensitive search across all 231 files. The 20 hits for registry are the model-alias registry, a different thing |
| Confirmed | The approval gate the package does implement fires on a tool’s name, from a fixed set — {"bash_", "edit_", "write_"} by default | Read in claude_loop_.py, where the gate is one name-membership test |
| Confirmed | In the unattended modes — the surface the paper’s cron-backbone argument rests on — that set is empty, so nothing is gated at all | Both headless and batch entry points pass an empty set, and their own comments say “no approval gate” |
| Not public | The engagements the architecture is drawn from. European enterprises across automotive, manufacturing, fast-moving consumer goods and healthcare, evaluated against Azure | Not checkable by anyone. The paper states plainly that individual client or employer engagements are not named or described, and asks to be judged on the architecture instead |
| Guesswork | That the distance between the proposal and the released artefact is specifically the governance layer | This report’s reading. It is an observation about which strings appear in one package, and says nothing about what exists unpublished |
| Guesswork | That the paper and the empirical study published five days later may be disagreeing about model tier rather than about schemas | This report’s reading again. Neither paper says it, and neither tests the other’s regime |
Timeline
A field, then a proposal, then a test of it Every row is a publication or a release date, and every one of them is public
- 24 Mar 2026Anthropic publishes its own harness-design work, arriving at a fixed three-agent architecture: planner, generator, evaluator
- 31 Mar 2026Terminal Agents Suffice for Enterprise Automation is posted — a terminal-and-filesystem agent matching or outperforming more complex architectures at a fraction of the cost. Revised to a third version on 5 August
- 5 Apr 2026Version 0.1.0 of the harness the paper will later name reaches PyPI, four and a half months ahead of the paper
- 10 Apr 2026Can Coding Agents Be General Agents? tests one against an open-core ERP system: simple tasks reliable, complex ones failing in four characteristic ways
- 14 Apr 2026Dive into Claude Code compares three harnesses and finds each answers the same design questions differently
- 7 May 2026Stop Comparing LLM Agents Without Disclosing the Harness argues, as a position paper, that harness configuration can outweigh model choice on long-horizon tasks between models of comparable capability
- 11 May 2026Beyond Autonomy argues agent frameworks prioritise autonomy over governability, and proposes risk-tiered review
- 13 May 2026A harness study on algorithm discovery finds that, at a fixed token budget, fewer algorithms thought about more deeply score higher than many shallow ones
- 18 May 2026Code as Agent Harness, a 42-author survey, names six open challenges, among them human oversight for safety-critical actions
- 12 Jun 2026HarnessX argues the opposite structural point — harnesses should evolve from their own traces, not stay static — and reports an average gain of 14.5% across five benchmarks
- 28 Jul 2026SHarD reaches the same thesis from a security angle: the unit to build and distribute is a hardened harness, engineered once and run unmodified everywhere
- 18 Aug 2026Version 0.2.58 of the harness ships. It is the sixty-ninth release, and the last before the paper
- 20 Aug 2026The paper is submitted at 23:44 UTC, naming that package as its reference implementation and offering the naming as its own disclosure
- 25 Aug 2026Five days later, an empirical study measures exactly this kind of wrapping and reports a mixed result — better in one of four model-task cells, worse in two
- 28 Aug 2026Version 0.2.72 ships — the eighty-third release, and the one read for this report
What has no date at all One item, and it is the one the four mechanisms are argued from
The engagements. The paper says the architecture is informed by work in large European enterprises across automotive, manufacturing, fast-moving consumer goods and healthcare, and that individual client or employer engagements are not named or described. It gives three worked examples — applying policy rules to a record in a line-of-business system, triaging inbound requests against a CRM, and a manufacturing check comparing a supplier’s certificate of conformance against a material norm — and none carries a date, a name or a measured outcome. That is a defensible position for anyone writing about client work, and it is stated rather than hidden, which is the part that matters. But it does mean the field evidence behind the four mechanisms cannot be placed in the chronology above, and cannot be checked by anyone at all.
- 3worked examples, none with a date, a client or a measured outcome
- 4sectors named, all in Europe, all evaluated against one cloud’s identity and policy primitives
The Argument
The claim, and the hinge it turns on Stated in the paper’s own words, then followed where it leads
The paper’s contribution is one sentence, and it is a good one: “When every solution shares one identical codebase, auditing N solutions collapses to reviewing N instruction files.” Custom code justified an upfront approval gate because looking back at it was expensive; if every team ships the same engine and differs only in two text files, that cost is gone, and so is the reason for the gate. Registration becomes a side effect of pushing code rather than a committee a team waits on. It is a real argument and it is well made.
It turns on a hinge the paper does not name. Reviewing N solutions collapses to N text files only once somebody has reviewed the shared codebase. That review is what the whole saving is borrowed against, and the paper never says who performs it, how often, or against what. The literature it cites makes the point sharper rather than softer: the position paper it leans on argues that the harness, not the model, is what determines behaviour, which is precisely why one unread harness under every solution is a larger unknown than N read ones, not a smaller one.
For the implementation it names, that review is legally constrained. The licence permits installing and running the package for internal business use, and forbids modifying, forking, decompiling, reverse engineering or creating derivative works without written permission. The Python source ships in the distribution, so the barrier is a legal one and not a technical one — but reverse engineering is where a security review of a dependency usually ends up, and here it is named as forbidden. None of this makes the architecture wrong. The paper says elsewhere that open-source harnesses already exist and are worth studying directly, so nothing about the design requires a closed one. The tension is with the artefact, not the idea.
- 83releases behind the URL offered as disclosure, and the paper names none of them
- 0mentions of
gatewayorentitlementin the shipped source
Three findings, and what the cited abstracts say The over-reach is in the paper’s own summary of the field; its related-work section states all three correctly
| As summarised | What the cited abstract says | Verdict |
|---|---|---|
| Harnesses suffice at the task level and outperform more elaborate architectures on enterprise work — on two citations | The first says terminal agents “match or outperform” more complex architectures “at a fraction of the cost”. The second runs no architecture comparison at all: it reports simple tasks succeeding and complex ones failing in four characteristic ways | Half supported. The first citation carries it, minus “match or” and minus the cost result, which is the stronger half of that paper’s finding. The second does not carry it |
| Harness choice accounts for most of the variance in agent benchmark results, more than model choice does — and the related-work section says the cited paper “establishes” it | It opens “This position paper argues”. Its claim is scoped to long-horizon tasks and to models of comparable frontier capability, and its wording is that harness-induced variance “can substantially exceed” model-induced variance | Two qualifiers dropped, and a position paper is not something that “establishes” a finding |
| The open gap between that finding and enterprise adoption is governance — on two citations | The first is squarely about governability and supports it. The second is a survey listing six open challenges, of which human oversight for safety-critical actions is one | Fair, mildly stretched. One item on a list of six is not the gap that list names |
The mechanism that is described but not shipped And the one that is shipped, which is the design the paper argues against
The paper’s risk mechanism is precise and unusual. Every tool call carries {params, risky}; the model sets risky from the parameters it has just constructed; a call that omits it is rejected before it reaches any backend. The paper is emphatic that this is the point: risk is call-scoped and self-declared, “never a static tier attached top-down to a tool or a tool group”. A flagged call then spawns a fresh instance of the same harness to judge it, so the context that composed the call is not the context grading it, and a human is asked only when that judge cannot clear it.
Searched case-insensitively, the word risky appears zero times in the 231 files of the published package. So do gateway, entitlement, config.yaml and live_skills. What the package does implement is an approval gate keyed on the tool’s name, tested against a set the user configures, defaulting to shell, edit and write — a static tier attached top-down to a tool. In the headless and batch entry points, which are the unattended surface the paper’s cron argument depends on, that set is passed empty and the source comments say so plainly.
This does not contradict the paper, and it would be dishonest to present it as if it did. The architecture puts authorization outside the harness on purpose, and the package is described as the harness. What it does show is where the line falls: everything a reader can install and check is the harness half, which was not the new part. The four services the governance argument rests on — gateway, solution registry, entitlement mapping, run-trigger endpoint — exist as prose and short snippets. The paper’s discussion section lists five limitations, and is candid about every one of them. This is not among the five.
- 231files searched, 101 of them Python source, none carrying the word
- 5limitations the paper lists and owns. This is not one of them
What Others Add
Four other pieces of work, and where each one lands Three of the four are not in the paper’s reference list; one of them postdates it by five days
25 Aug 2026 · not cited
The controlled study, published five days later
- Wrapping an agent in a deterministic layer improved reproducibility in one of four model-task cells, degraded it in two, and did nothing in the fourth
- The trace-level diagnosis: once tool and state sequences are already consistent, an unconstrained free-text planning step becomes the dominant source of variance
- Its fix is validating the plan against a fixed schema before any tool runs — which is the thing this paper calls a net negative, arguing that params should be left open
- The catch, and it is a real one: its models are a 7B and a 27B open-weight model on synthetic tasks, while this paper’s entire bet is on frontier capability. Neither tests the other’s regime, so this may be a disagreement about model tier rather than about schemas — a reading neither paper offers
12 Jun 2026 · not cited
The opposite structural bet
- Its complaint is that harnesses “remain largely hand-crafted and static”, and that the traces a run produces are rarely fed back into improving them
- It reports an average gain of 14.5%, up to 44.0%, across five benchmarks, with the gains largest where the baseline was weakest
- That is a direct trade against this paper’s central move. One harness, identical everywhere, is exactly what makes the fleet auditable; adapting it per deployment is what buys the performance
- Both can be right at once. The paper simply does not name the trade, and a reader weighing the architecture should
28 Jul 2026 · cited
The security paper it agrees with, and the two controls it drops
- It names three controls a distributable harness can carry: operating-system sandboxing, skill scanning, and tool restriction
- This paper develops the third, and says so. Its own discussion concedes that a container is a coarser guarantee than the sandboxing SHarD tested, and that skill scanning has no equivalent at all
- It goes further, and this is the most useful admission in the whole discussion: a skill fetched live at every run widens that gap rather than narrowing it, because there is no rebuild step where a scan could be inserted and no image to diff against a known-good one
24 Mar 2026 · not cited
Anthropic’s own harness work fixes the graph
- The paper’s second design decision is that there is no graph fixed at design time: a single instance calls the model repeatedly and composes its own orchestration at runtime
- Anthropic’s published harness work of March 2026 arrives at the opposite: a three-agent structure, decided in advance, which it credits for multi-hour autonomous runs
- The two are not straightforwardly opposed. The paper’s judge is an evaluator, spawned per flagged call, and the paper builds openly on Anthropic’s 2025 multi-agent write-up, which it cites accurately
- Still, a paper titled for applying one lab’s primitives makes a structural choice that lab’s own published harness work does not
Conclusion
What holds, what is still a proposal, and what to hold loosely A paper without a benchmark is not a paper without evidence — but the evidence has to be named for what it is
What holds up, and it is more than most
The citation record is clean. All eight arXiv references resolve and every author list matches the record exactly and in order, one of them 42 names long — that is not the norm for a single-author preprint. The reference implementation is real, current, and matches the paper in detail a reader can check: the status file tracking each spawned child, the socket-or-inbox messaging between them, the browser interface, the structured status line ending every run, and a model registry whose default is the very model the configuration example names.
What is still a proposal
The four services the governance argument rests on. A gateway that decides entitlement, a registry that records what was deployed and by whom, an entitlement mapping in version-controlled YAML, and a run-trigger endpoint: each is described carefully, each is illustrated with a short snippet, and none of them is in anything a reader can install. That is a legitimate way to publish an architecture — but the harness half was already the well-trodden part, so what a reader can check is the part that was not new. The paper is candid about having run no benchmark and lists five limitations. This report’s reading, not the paper’s: the distance between the proposal and the artefact is a sixth.
The claim worth weighing carefully
“Auditing N solutions collapses to reviewing N instruction files” is only as good as the single review of the shared codebase it is borrowed against, and the paper never says who does that review. For the implementation it names, the licence permits installing and running but forbids modifying, forking and reverse engineering without written permission — and reverse engineering is where a serious review of a dependency tends to end up. The source ships, so the barrier is legal rather than technical, and nothing in the architecture requires a closed harness. But an enterprise adopting this pattern is choosing which harness sits under every solution it will ever run, and that choice deserves the same scrutiny the pattern promises to make cheap everywhere else.
Hold this loosely
Everything counted here is a snapshot of one version, read eight days after the paper was submitted. Nothing in the release history suggests a governance layer was removed — it was in none of the releases sampled — but a snapshot is what it is, and the numbers above are dated for that reason. The field evidence is uncheckable by design: the sectors are named, the clients are not, and the author says so rather than dressing it up, which is the honest version of an unverifiable claim. And a paper that grades its own confidence carefully in its discussion section, as this one does, has earned a reading that separates what it argues from what it has shipped.