Tratopedia
繁中
Settings

Text size

Language

Theme

High contrast

Version

v1.81.0

The release this page was built from. It is what the service worker caches under.

AI · LLM Engineering · TopicTratopedia · revised 16 Aug 2026

Standing topic · five names, one moving boundary

LLM engineering

LLM engineering is the umbrella for building systems around a language model rather than training one. It has five names. Prompt engineering was a dictionary headword by 2023; the other four — context, harness, loop and graph engineering — all arrived between June 2025 and July 2026, inside 380 days. The field presents them as a ladder, each layer containing the last. Five separate investigations, one per term, do not support that. No two of the five moved the same way: one narrowed, one inverted, one swelled, one never converged, and one was coined twice a fortnight apart by people who appear not to have known of each other. In every case this project could date — four out of four — the practice was published before the word, by 71 to 182 days. Not one of those four coinages happened in a peer-reviewed venue. And nothing anywhere in the material measures whether any layer works better than the one below it. What is agreed, across all five arguments about names, is the practice itself — and it has barely moved.

The five terms, and how each one moved each has its own article; this is what they show side by side

  • 5named layers, and no two of them moved the same way
  • 380days from the first datable coinage to the last
  • 4 / 4times the practice was published before the word
  • 0 / 4datable coinages introduced in a peer-reviewed venue
  • 1comparison in the whole corpus with a controlled shape
  • Prompt engineering 提示工程

    It narrowed

    coinage date not established · a headword by 2023

    • Task selector, then teaching by example, then process shaping
    • Then a job title — and then a component inside the layer above it
    • The job ended. The measured sensitivity to wording never did
  • Context engineering 情境工程

    It inverted

    19 Jun 2025 · inverted in 102 days

    • Coined as “providing all the context”
    • Formalised as “the smallest possible set of tokens”
    • Sufficiency became scarcity, and once-per-task became once-per-step
  • Harness engineering 機具工程

    It swelled

    5 Feb 2026 · redefined 33 days later

    • Coined for a habit: an instructions file and a couple of scripts
    • Thirty-three days later: Agent = Model + Harness
    • “If you're not the model, you're the harness”
  • Loop engineering 迴圈工程

    It never converged

    2–7 Jun 2026 · six framings in 29 days

    • Neither of the two people who started it ever defined it
    • Six accounts of five different objects — a person, a goal, a program, two machines, an organisation
    • One vendor conceded in print that nobody agrees what a loop is
  • Graph engineering 圖結構工程

    It was coined twice

    4 Jul 2026 in earnest · 18 Jul 2026 as a joke

    • The careful, sourced essay sank. The twelve-word joke drew 3.1M views
    • A survey had already named the practice in January: flow engineering
    • The technical content is the least disputed of the five

How much of this is established every claim on this page, graded — sorted by standing, never mixed

StandingWhat is claimedWhere it comes from
ConfirmedThe four coinage dates: 19 Jun 2025, 5 Feb 2026, 2–7 Jun 2026, 4 Jul 2026Each read at source; the two X timestamps re-derived from their post IDs for this page
ConfirmedThe four movements: 102 days to invert, 33 to swell, 29 to fragment, 14 between the two coinagesRecomputed from the dated documents in this topic's research record, not copied from the articles
ConfirmedA January 2026 survey already called the graph practice flow engineeringarXiv:2601.12560 §5.1, read in full. The essay that coined “graph engineering” cites this paper for something else
ConfirmedWikipedia has an article on the agent harness and none on context engineeringBoth checked on 15 Aug 2026; the second returned HTTP 404. Loop and graph engineering were not checked
Confirmed, not verifiable hereCherny's “my job is to write loops” — the sentence most often given as loop engineering's originSpoken at an event on 2 Jun 2026 and reaching this project only through a third party's write-up. Never heard at source
Confirmed, not verifiable hereAnthropic's two definitions — the augmented LLM, and “the smallest possible set of tokens”anthropic.com refuses this session; the text was supplied by djTratoh and is summarised in input/summary/
UnconfirmedWhen “prompt engineering” was first said, and by whomNever established. It is therefore excluded from the four-out-of-four pattern and from the 380-day span, rather than assumed to fit them
This project's readingThat each name arrives when a particular unit of work stops being the one a person controlsAn inference drawn across the five records. No source states it, and it is marked as inference wherever it appears
This project's readingThat four gaps between practice and word amount to a pattern rather than a coincidenceFour instances with no counterexample found. Four is not many, and the page says so
Not publicThe contents of the X Article that announced loop engineering's deathBehind X Premium. Title, byline and timestamp only — this page characterises none of its contents

How the vocabulary arrived the mechanisms are 2018–2022; the names are 2025–2026; there is very little in between

  1. 2018decaNLP casts ten tasks as questions over a context. The prompt is a task selector
  2. 2020GPT-3: 175B parameters, few-shot sometimes competitive with fine-tuning. The prompt now carries the examples
  3. Jan 2022Chain-of-thought: eight exemplars beat a fine-tuned model with a verifier. The prompt shapes the process — ten months before ChatGPT
  4. Oct 2022ReAct: reason, act, observe. The inner loop exists as a paper, four years before anyone names loop engineering
  5. 19 Dec 2024Anthropic publishes the augmented LLM — retrieval, tools, memory. This is context engineering, six months before it has a name
  6. 19 Jun 2025Context engineering is named, on X, as “providing all the context”. 182 days after the practice was published
  7. 25 Jun 2025Six days later a second definition: “just the right information for the next step”. Already narrower, and already per-turn
  8. 14 Jul 2025Chroma measures context rot across 18 models: they do not use their context uniformly, and get less reliable as it grows
  9. 29 Sep 2025“The smallest possible set of high-signal tokens.” The term has inverted in 102 days
  10. 26 Nov 2025Anthropic publishes the whole harness practice — an initializer, a feature file, a progress file. The word “harness” is used as an ordinary noun
  11. 18 Jan 2026A survey names the graph practice: it “is often described as flow engineering”. Nobody in the later material mentions this
  12. 23 Jan 2026OpenAI decomposes the inner loop in public, and calls its agent “our agent (or harness)” — the noun needing no defence, two weeks before it becomes a discipline
  13. 5 Feb 2026Harness engineering is coined, on a personal blog, for an instructions file and two scripts. Its author says he knows of no existing term
  14. 10 Mar 2026Agent = Model + Harness. Thirty-three days, and the word now means everything that is not the model
  15. 13 Apr 2026A paper formalises the loop as a scheduler whose ready set is exactly one — and calls its own proposal a Structured Graph Harness, three months before “graph engineering” is said
  16. 2 Jun 2026“I don't prompt Claude anymore. My job is to write loops.” A report of one person's practice, at an event. 130 days after the inner loop was published
  17. 7 Jun 2026An X post turns that report into an instruction — “you should be designing loops that prompt your agents” — and it travels. Neither post defines a loop
  18. Jun – 1 Jul 2026Six framings in 29 days, of five different objects: a person's habits, a goal, a program, two nested machines, a taxonomy, an organisation whose outer loop runs for weeks and contains users
  19. 4 Jul 2026Graph engineering is coined in earnest — seven minutes, six citations, three commitments. It sinks without trace
  20. 18 Jul 2026Twelve words on X, as a joke, at 00:34 UTC — 3.1M views. An obituary for loop engineering follows 4h35m later, behind a paywall. The joke is the version that travels
  21. 27–28 Jul 2026A Taiwanese trade title draws the five-layer ladder, dating every rung — and in the text beside it writes 目前還沒有定論, there is no settled conclusion
  22. 13 Aug 2026Wikipedia carries an entry for the agent harness, its etymology marked contested. There is still no article for context engineering, fourteen months after it was named

The field's account of itself: a ladder asserted by the least hedged sources, and by nobody who did the naming

The standard account is a stack: prompt engineering inside context engineering, both inside the harness, the loop above that, the graph above the loop. It is a good story and it is drawn far more confidently than it is argued. The clearest statement of it is a diagram, which dates every layer and fixes the nesting; the article printed beside that diagram, by the same publisher on the same day, says in as many words that there is no settled conclusion about the order — some put the loop above the harness, and some leave the harness out altogether. One of those who leaves it out is the writer who coined graph engineering: his ladder has four rungs, and he omits the harness while leaning hardest on a paper whose proposal is called a Structured Graph Harness. The pattern is consistent enough to be worth naming. Confidence in this vocabulary rises with distance from the person who invented it.

  • 5 vs 4rungs, depending on whether the harness is a layer or a container
  • 4 / 4namers who hedged: “no existing term”, “I'm skeptical”, “should matter less over time”, and a joke

What the records show instead: a boundary moving outwards the reading is this project's; the movements underneath it are not

Read in date order, the five names do not stack. Each one arrives at the moment a particular unit of work stops being the unit a person can control. While the unit was one string, the discipline was prompt engineering. When it became one window, context engineering. When the window stopped being the edge of the system, everything outside the model — the harness. Then the repetition — the loop. Then what happens between windows — the graph. That reading is an inference and is marked as one; no source states it. What is not an inference is the consequence: five terms, and five different failure modes. Prompt engineering narrowed until it was a component. Context engineering inverted. Harness engineering swelled. Loop engineering fragmented into six accounts of five different objects. Graph engineering was coined twice, and the coinage that won was not meant seriously. Not one of the five simply came to mean one thing. A vocabulary that behaves this way is not a ladder being climbed. It is a field naming a boundary it keeps having to move.

  • 182 → 71days between practice and word — the gap narrows as the pace picks up
  • 41days apart, one person's two posts propelled two of the five names. Neither defines anything

What is actually agreed assembled across all five terms, and disputed by none of them

  1. 1Write at the right altitudeSpecific enough to guide, general enough to survive the next case. Not a list of every situation you have met
  2. 2Keep the tool set minimal and non-overlappingTwo tools that do nearly the same thing are a decision the model now has to get right, for no gain
  3. 3One job per nodeReached independently by three sources across two different terms. “A good node is boring”
  4. 4Verify from outsideThe system exists to disagree with the model's “I am done”. Asking a model to review its own answer is not external verification
  5. 5Checkpoint the state, not the transcriptState with a schema, written at every step. If memory is “whatever survived compaction”, failure at step 40 means starting again
  6. 6Put the budget in the stateThree sources, three terms, one sentence: if you cannot stop the system at a spend threshold, it is not autonomous — it is running up a bill

The evidence, and where it runs out the weakest part of the subject, and the field says so about itself

QuestionWhat existsStanding
Does the harness matter?One comparison with the right shape — model fixed, harness varied, on a third-party leaderboard. The same frontier model scores far apart in two harnessesThe strongest thing in the whole corpus, and it covers one term of five
Does a loop beat prompting?Nothing. No controlled comparison was found in any sourceAn absence found by looking for it, not a gap in the search
Does a graph beat a loop?Nothing. The one rigorous paper says of itself: “not a production implementation or empirical results”, under an explicit fairness disclaimerTwo terms running now rest on no measurement at all
Is the paper trail reliable?Not yet. An encyclopaedia entry mis-dates its own primary source by a year; the one survey behind the graph case reports 42, 41, 49 and 68 in four places for a study stated as 70Both checked against the documents. Check dates and URLs yourself; this material is young

Two facts about where this vocabulary comes from are worth more than any of the advocacy. The first: of the four coinages this project could date, none happened in a peer-reviewed venue — two X posts, a personal blog, a personal site and a conference talk between them. The papers are real and they arrive afterwards, adopting a word the field has already chosen. The second: the people who did the naming are the ones least sure about it. The man who coined harness engineering said he knew of no existing term and did not press the point. The man who named loop engineering wrote “its still early, I'm skeptical” in the same paragraph. The man whose definition of the harness everyone else restates argues that harnesses should matter less over time. And the post that made graph engineering famous was a joke, which a named analyst pointed out four days later — while insisting, correctly, that the joke worked because it pointed at something real. That last distinction is the one to keep. The naming is noise. The practice under it is not.

So what what to take from four names in 380 days

The practice comes first. Every time.

Four terms, four datable gaps, no counterexample: 182 days for context engineering, 71 for harness, 130 for loop, 167 for graph. In the graph case the practice was not merely being done — it had already been named, in a January survey, as flow engineering. If a new term appears, the useful question is not what it means but what has already been shipping under another name for the last six months.

This is a vocabulary, not a technology.

The mechanisms are 2018 to 2022 and they are published and peer-reviewed. The names are 2025 to 2026, and not one of the four datable coinages happened in a peer-reviewed venue. Almost nothing sits in between. That does not make the terms worthless — a field needs words — but it does mean a new name on this ladder is evidence about the conversation, not about the systems.

Trust the namers over the repeaters.

Every one of the four hedged — “I know of no existing term”, “I'm skeptical”, “harnesses should matter less over time”, and one outright joke. Downstream, the hedges vanish: the confident five-rung ladder is drawn by a publisher whose own article says there is no settled conclusion. Where a claim about this vocabulary sounds certain, check how far it is from whoever invented the word.

And the part that actually holds.

Right altitude, minimal tools, one job per node, external verification, checkpointed state, a budget you can stop the system at. Six practices, assembled from five separate arguments about what to call the layer they belong to, and contradicted by none of them. That is the durable part of this subject, and it would still be worth doing if every one of the five names were withdrawn tomorrow. It is also, so far, the part nobody has measured.

You are reading v0001, published 2026-08-16. It has been superseded — the current version is v0002.

Articles on this topic

In reading order. A series is kept together and starts at its first part.

  • 1/5 · First of five · the innermost layer

    The Job Title Died. The Sensitivity Did Not.

    Prompt engineering is writing the instruction so the model does what you meant. It got a dictionary entry, a job title, and then an obituary: the trade press called the role obsolete within two years of inventing it.…

  • 2/5 · Second of five · the layer around the prompt

    “All the Context” Lasted Six Days

    On 19 June 2025 Shopify's CEO said context engineering was the art of providing all the context for a task to be plausibly solvable. Six days later Andrej Karpathy endorsed the term and quietly changed it: filling the…

  • 3/5 · Third of five · the outermost layer

    Coined for a Text File. A Month Later It Meant Everything.

    On 5 February 2026 Mitchell Hashimoto needed a name for a habit — fix the agent's environment every time it errs — and, finding none, called it harness engineering. What he meant by it was an instruction file and a…

  • 4/5 · Fourth of five · the one that never settled

    Everyone Agrees You Should Write Loops. Nobody Agrees What One Is.

    In the first week of June 2026 two well-known engineers said, four days apart, that you should stop prompting coding agents and write loops instead. Neither said what a loop was. In the three weeks that followed, six…

  • 5/5 · Fifth of five · the one that arrived twice

    One Term, Coined Twice, Two Weeks Apart. The Joke Won.

    On 4 July 2026 Josh C. Simmons published a careful, sourced essay coining graph engineering. It sank. Fourteen days later, at 00:34 UTC on 18 July, Peter Steinberger asked twelve words on X — Are we still talking loops…

  • · Source review · five names, checked one at a time

    The Picture Is Surer Than the Text

    BusinessNext introduced five kinds of “AI engineering” in a graphic and an article: prompt, context, harness, loop, graph. Both say the same thing about how they relate — not new replacing old, but each one wrapping…

Versions

This summary is rewritten as the corpus grows. Every version stays published at its own address.

  1. v0002 current
  2. v0001 superseded

The current version is also at latest/.