Tratopedia
Side N繁中
Settings

Text size

Language

Theme

High contrast

Version

v1.177.0

The release this page was built from. It is what the service worker caches under.

AI Agents · Governance · Software Engineering · Benchmarking · Research Briefpaper of 20 Aug 2026 · package read at version 0.2.72, 28 Aug

arXiv:2608.20622 · a proposal, and the package it names

The Harness Ships, the Governance Doesn’t

On 20 August 2026 George Salapa posted a 21-page architecture proposal to arXiv: give a large enterprise one coding-agent harness, run it unmodified behind every surface, and the whole fleet becomes auditable, because reviewing a solution stops being a code review and becomes reading one instructions file. The paper names its reference implementation, which is a real Python package with 83 releases behind it. Reading that package is what this report does. The harness the paper describes is there in detail. The four services its governance argument rests on are not — and its licence forbids the forking on which the deployment model turns.

The figures Each one counted from a record anyone can fetch: the arXiv API, the PyPI API, and the package’s own source distribution

  • 83releases behind the one URL the paper offers as its disclosure. It names none of them
  • 0times risky — the flag the paper’s whole risk mechanism turns on — appears anywhere in the shipped package
  • 8/8arXiv references that resolve, every author list matching the record in order — one of them 42 names long
  • 0benchmarks in the paper, which says so itself: “No benchmark accompanies these claims”

The record, graded Sorted by how well each stands up, not by how interesting it is

StandingWhatHow it was checked
ConfirmedA 21-page architecture proposal was submitted to arXiv at 23:44 UTC on 20 August 2026, in cs.AI and cs.SE, by one author, with nine references and three figuresThe arXiv API record and the abstract page agree with the document
ConfirmedAll eight arXiv references exist, and every author list in the reference section matches the record exactly and in order — including a 42-author entryEach identifier queried against the arXiv API and compared name by name
ConfirmedThe reference implementation it names is real and current: 83 releases from 5 April to 28 August 2026, Python source included in the distributionThe PyPI JSON record, then the source distribution downloaded and listed
ConfirmedIts licence permits installing and running it for internal business use, and forbids copying, modifying, forking, decompiling, reverse engineering or creating derivative works without written permissionThe LICENSE file inside the distribution; the PyPI classifier reads License :: Other/Proprietary License
ConfirmedThe paper names no version. Sixty-nine releases preceded it and fourteen have followed, so the URL it cites has not meant the same thing twiceUpload timestamps for all 83 releases, split at the paper’s submission time
ConfirmedNothing in the package names the governance layer: risky, gateway, entitlement, config.yaml and live_skills each appear zero timesCase-insensitive search across all 231 files. The 20 hits for registry are the model-alias registry, a different thing
ConfirmedThe approval gate the package does implement fires on a tool’s name, from a fixed set — {"bash_", "edit_", "write_"} by defaultRead in claude_loop_.py, where the gate is one name-membership test
ConfirmedIn the unattended modes — the surface the paper’s cron-backbone argument rests on — that set is empty, so nothing is gated at allBoth headless and batch entry points pass an empty set, and their own comments say “no approval gate”
Not publicThe engagements the architecture is drawn from. European enterprises across automotive, manufacturing, fast-moving consumer goods and healthcare, evaluated against AzureNot checkable by anyone. The paper states plainly that individual client or employer engagements are not named or described, and asks to be judged on the architecture instead
GuessworkThat the distance between the proposal and the released artefact is specifically the governance layerThis report’s reading. It is an observation about which strings appear in one package, and says nothing about what exists unpublished
GuessworkThat the paper and the empirical study published five days later may be disagreeing about model tier rather than about schemasThis report’s reading again. Neither paper says it, and neither tests the other’s regime

A field, then a proposal, then a test of it Every row is a publication or a release date, and every one of them is public

  1. 24 Mar 2026Anthropic publishes its own harness-design work, arriving at a fixed three-agent architecture: planner, generator, evaluator
  2. 31 Mar 2026Terminal Agents Suffice for Enterprise Automation is posted — a terminal-and-filesystem agent matching or outperforming more complex architectures at a fraction of the cost. Revised to a third version on 5 August
  3. 5 Apr 2026Version 0.1.0 of the harness the paper will later name reaches PyPI, four and a half months ahead of the paper
  4. 10 Apr 2026Can Coding Agents Be General Agents? tests one against an open-core ERP system: simple tasks reliable, complex ones failing in four characteristic ways
  5. 14 Apr 2026Dive into Claude Code compares three harnesses and finds each answers the same design questions differently
  6. 7 May 2026Stop Comparing LLM Agents Without Disclosing the Harness argues, as a position paper, that harness configuration can outweigh model choice on long-horizon tasks between models of comparable capability
  7. 11 May 2026Beyond Autonomy argues agent frameworks prioritise autonomy over governability, and proposes risk-tiered review
  8. 13 May 2026A harness study on algorithm discovery finds that, at a fixed token budget, fewer algorithms thought about more deeply score higher than many shallow ones
  9. 18 May 2026Code as Agent Harness, a 42-author survey, names six open challenges, among them human oversight for safety-critical actions
  10. 12 Jun 2026HarnessX argues the opposite structural point — harnesses should evolve from their own traces, not stay static — and reports an average gain of 14.5% across five benchmarks
  11. 28 Jul 2026SHarD reaches the same thesis from a security angle: the unit to build and distribute is a hardened harness, engineered once and run unmodified everywhere
  12. 18 Aug 2026Version 0.2.58 of the harness ships. It is the sixty-ninth release, and the last before the paper
  13. 20 Aug 2026The paper is submitted at 23:44 UTC, naming that package as its reference implementation and offering the naming as its own disclosure
  14. 25 Aug 2026Five days later, an empirical study measures exactly this kind of wrapping and reports a mixed result — better in one of four model-task cells, worse in two
  15. 28 Aug 2026Version 0.2.72 ships — the eighty-third release, and the one read for this report

What has no date at all One item, and it is the one the four mechanisms are argued from

The engagements. The paper says the architecture is informed by work in large European enterprises across automotive, manufacturing, fast-moving consumer goods and healthcare, and that individual client or employer engagements are not named or described. It gives three worked examples — applying policy rules to a record in a line-of-business system, triaging inbound requests against a CRM, and a manufacturing check comparing a supplier’s certificate of conformance against a material norm — and none carries a date, a name or a measured outcome. That is a defensible position for anyone writing about client work, and it is stated rather than hidden, which is the part that matters. But it does mean the field evidence behind the four mechanisms cannot be placed in the chronology above, and cannot be checked by anyone at all.

  • 3worked examples, none with a date, a client or a measured outcome
  • 4sectors named, all in Europe, all evaluated against one cloud’s identity and policy primitives

The claim, and the hinge it turns on Stated in the paper’s own words, then followed where it leads

The paper’s contribution is one sentence, and it is a good one: “When every solution shares one identical codebase, auditing N solutions collapses to reviewing N instruction files.” Custom code justified an upfront approval gate because looking back at it was expensive; if every team ships the same engine and differs only in two text files, that cost is gone, and so is the reason for the gate. Registration becomes a side effect of pushing code rather than a committee a team waits on. It is a real argument and it is well made.

It turns on a hinge the paper does not name. Reviewing N solutions collapses to N text files only once somebody has reviewed the shared codebase. That review is what the whole saving is borrowed against, and the paper never says who performs it, how often, or against what. The literature it cites makes the point sharper rather than softer: the position paper it leans on argues that the harness, not the model, is what determines behaviour, which is precisely why one unread harness under every solution is a larger unknown than N read ones, not a smaller one.

For the implementation it names, that review is legally constrained. The licence permits installing and running the package for internal business use, and forbids modifying, forking, decompiling, reverse engineering or creating derivative works without written permission. The Python source ships in the distribution, so the barrier is a legal one and not a technical one — but reverse engineering is where a security review of a dependency usually ends up, and here it is named as forbidden. None of this makes the architecture wrong. The paper says elsewhere that open-source harnesses already exist and are worth studying directly, so nothing about the design requires a closed one. The tension is with the artefact, not the idea.

  • 83releases behind the URL offered as disclosure, and the paper names none of them
  • 0mentions of gateway or entitlement in the shipped source

Three findings, and what the cited abstracts say The over-reach is in the paper’s own summary of the field; its related-work section states all three correctly

As summarisedWhat the cited abstract saysVerdict
Harnesses suffice at the task level and outperform more elaborate architectures on enterprise work — on two citationsThe first says terminal agents “match or outperform” more complex architectures “at a fraction of the cost”. The second runs no architecture comparison at all: it reports simple tasks succeeding and complex ones failing in four characteristic waysHalf supported. The first citation carries it, minus “match or” and minus the cost result, which is the stronger half of that paper’s finding. The second does not carry it
Harness choice accounts for most of the variance in agent benchmark results, more than model choice does — and the related-work section says the cited paper “establishes” itIt opens “This position paper argues”. Its claim is scoped to long-horizon tasks and to models of comparable frontier capability, and its wording is that harness-induced variance “can substantially exceed” model-induced varianceTwo qualifiers dropped, and a position paper is not something that “establishes” a finding
The open gap between that finding and enterprise adoption is governance — on two citationsThe first is squarely about governability and supports it. The second is a survey listing six open challenges, of which human oversight for safety-critical actions is oneFair, mildly stretched. One item on a list of six is not the gap that list names

The mechanism that is described but not shipped And the one that is shipped, which is the design the paper argues against

The paper’s risk mechanism is precise and unusual. Every tool call carries {params, risky}; the model sets risky from the parameters it has just constructed; a call that omits it is rejected before it reaches any backend. The paper is emphatic that this is the point: risk is call-scoped and self-declared, “never a static tier attached top-down to a tool or a tool group”. A flagged call then spawns a fresh instance of the same harness to judge it, so the context that composed the call is not the context grading it, and a human is asked only when that judge cannot clear it.

Searched case-insensitively, the word risky appears zero times in the 231 files of the published package. So do gateway, entitlement, config.yaml and live_skills. What the package does implement is an approval gate keyed on the tool’s name, tested against a set the user configures, defaulting to shell, edit and write — a static tier attached top-down to a tool. In the headless and batch entry points, which are the unattended surface the paper’s cron argument depends on, that set is passed empty and the source comments say so plainly.

This does not contradict the paper, and it would be dishonest to present it as if it did. The architecture puts authorization outside the harness on purpose, and the package is described as the harness. What it does show is where the line falls: everything a reader can install and check is the harness half, which was not the new part. The four services the governance argument rests on — gateway, solution registry, entitlement mapping, run-trigger endpoint — exist as prose and short snippets. The paper’s discussion section lists five limitations, and is candid about every one of them. This is not among the five.

  • 231files searched, 101 of them Python source, none carrying the word
  • 5limitations the paper lists and owns. This is not one of them

Four other pieces of work, and where each one lands Three of the four are not in the paper’s reference list; one of them postdates it by five days

  • 25 Aug 2026 · not cited

    The controlled study, published five days later

    Saransh Dhage, on deterministic execution constraints

    • Wrapping an agent in a deterministic layer improved reproducibility in one of four model-task cells, degraded it in two, and did nothing in the fourth
    • The trace-level diagnosis: once tool and state sequences are already consistent, an unconstrained free-text planning step becomes the dominant source of variance
    • Its fix is validating the plan against a fixed schema before any tool runs — which is the thing this paper calls a net negative, arguing that params should be left open
    • The catch, and it is a real one: its models are a 7B and a 27B open-weight model on synthetic tasks, while this paper’s entire bet is on frontier capability. Neither tests the other’s regime, so this may be a disagreement about model tier rather than about schemas — a reading neither paper offers
  • 12 Jun 2026 · not cited

    The opposite structural bet

    HarnessX, on harnesses that evolve from their own traces

    • Its complaint is that harnesses “remain largely hand-crafted and static”, and that the traces a run produces are rarely fed back into improving them
    • It reports an average gain of 14.5%, up to 44.0%, across five benchmarks, with the gains largest where the baseline was weakest
    • That is a direct trade against this paper’s central move. One harness, identical everywhere, is exactly what makes the fleet auditable; adapting it per deployment is what buys the performance
    • Both can be right at once. The paper simply does not name the trade, and a reader weighing the architecture should
  • 28 Jul 2026 · cited

    The security paper it agrees with, and the two controls it drops

    SHarD, reaching the same thesis from a different direction

    • It names three controls a distributable harness can carry: operating-system sandboxing, skill scanning, and tool restriction
    • This paper develops the third, and says so. Its own discussion concedes that a container is a coarser guarantee than the sandboxing SHarD tested, and that skill scanning has no equivalent at all
    • It goes further, and this is the most useful admission in the whole discussion: a skill fetched live at every run widens that gap rather than narrowing it, because there is no rebuild step where a scan could be inserted and no image to diff against a known-good one
  • 24 Mar 2026 · not cited

    Anthropic’s own harness work fixes the graph

    A planner, a generator and an evaluator, chosen at design time

    • The paper’s second design decision is that there is no graph fixed at design time: a single instance calls the model repeatedly and composes its own orchestration at runtime
    • Anthropic’s published harness work of March 2026 arrives at the opposite: a three-agent structure, decided in advance, which it credits for multi-hour autonomous runs
    • The two are not straightforwardly opposed. The paper’s judge is an evaluator, spawned per flagged call, and the paper builds openly on Anthropic’s 2025 multi-agent write-up, which it cites accurately
    • Still, a paper titled for applying one lab’s primitives makes a structural choice that lab’s own published harness work does not

What holds, what is still a proposal, and what to hold loosely A paper without a benchmark is not a paper without evidence — but the evidence has to be named for what it is

What holds up, and it is more than most

The citation record is clean. All eight arXiv references resolve and every author list matches the record exactly and in order, one of them 42 names long — that is not the norm for a single-author preprint. The reference implementation is real, current, and matches the paper in detail a reader can check: the status file tracking each spawned child, the socket-or-inbox messaging between them, the browser interface, the structured status line ending every run, and a model registry whose default is the very model the configuration example names.

What is still a proposal

The four services the governance argument rests on. A gateway that decides entitlement, a registry that records what was deployed and by whom, an entitlement mapping in version-controlled YAML, and a run-trigger endpoint: each is described carefully, each is illustrated with a short snippet, and none of them is in anything a reader can install. That is a legitimate way to publish an architecture — but the harness half was already the well-trodden part, so what a reader can check is the part that was not new. The paper is candid about having run no benchmark and lists five limitations. This report’s reading, not the paper’s: the distance between the proposal and the artefact is a sixth.

The claim worth weighing carefully

“Auditing N solutions collapses to reviewing N instruction files” is only as good as the single review of the shared codebase it is borrowed against, and the paper never says who does that review. For the implementation it names, the licence permits installing and running but forbids modifying, forking and reverse engineering without written permission — and reverse engineering is where a serious review of a dependency tends to end up. The source ships, so the barrier is legal rather than technical, and nothing in the architecture requires a closed harness. But an enterprise adopting this pattern is choosing which harness sits under every solution it will ever run, and that choice deserves the same scrutiny the pattern promises to make cheap everywhere else.

Hold this loosely

Everything counted here is a snapshot of one version, read eight days after the paper was submitted. Nothing in the release history suggests a governance layer was removed — it was in none of the releases sampled — but a snapshot is what it is, and the numbers above are dated for that reason. The field evidence is uncheckable by design: the sectors are named, the clients are not, and the author says so rather than dressing it up, which is the honest version of an unverifiable claim. And a paper that grades its own confidence carefully in its discussion section, as this one does, has earned a reading that separates what it argues from what it has shipped.

Sources — Salapa, Applying Anthropic Primitives at Large Enterprises: Harness Paradigm for Knowledge Work, arXiv:2608.20622v1, 20 Aug 2026. Bibliographic records, author lists and abstracts for arXiv:2604.00073 (v3), 2604.13107, 2604.14228 (v2), 2605.10223, 2605.15221, 2605.18747, 2605.23950, 2606.14249, 2607.25890 and 2608.26197, read via the arXiv API. micro-cc package metadata and release history, PyPI JSON API; licence text, file inventory and source read from micro_cc-0.2.72.tar.gz as published on PyPI. Anthropic, How we built our multi-agent research system, 13 Jun 2025, and Harness design for long-running application development, 24 Mar 2026. Microsoft Learn, driveItem: delta, Microsoft Graph v1.0 reference. Every source fetched 30 Aug 2026. Only the abstracts of the ten cited and comparison papers were read, via the arXiv API; every statement here about them is a statement about their abstracts.

Versions

This document is rewritten when what it says has to change. Every version stays published at its own address.

  1. v0001 current

The current version is also at latest/.