What It Is
AI · LLM Engineering · TopicTratopedia · revised 16 Aug 2026
Standing topic · five names, one moving boundary
LLM engineering
LLM engineering is the umbrella for building systems around a language model rather than training one. It has five names. Prompt engineering was a dictionary headword by 2023; the other four — context, harness, loop and graph engineering — all arrived between June 2025 and July 2026, inside 380 days. The field presents them as a ladder, each layer containing the last. Five separate investigations, one per term, do not support that. No two of the five moved the same way: one narrowed, one inverted, one swelled, one never converged, and one was coined twice a fortnight apart by people who appear not to have known of each other. In every case this project could date — four out of four — the practice was published before the word, by 71 to 182 days. Not one of those four coinages happened in a peer-reviewed venue. And nothing anywhere in the material measures whether any layer works better than the one below it. What is agreed, across all five arguments about names, is the practice itself — and it has barely moved.
The five terms, and how each one moved each has its own article; this is what they show side by side
- 5named layers, and no two of them moved the same way
- 380days from the first datable coinage to the last
- 4 / 4times the practice was published before the word
- 0 / 4datable coinages introduced in a peer-reviewed venue
- 1comparison in the whole corpus with a controlled shape
Prompt engineering 提示工程
It narrowed
- Task selector, then teaching by example, then process shaping
- Then a job title — and then a component inside the layer above it
- The job ended. The measured sensitivity to wording never did
Context engineering 情境工程
It inverted
- Coined as “providing all the context”
- Formalised as “the smallest possible set of tokens”
- Sufficiency became scarcity, and once-per-task became once-per-step
Harness engineering 機具工程
It swelled
- Coined for a habit: an instructions file and a couple of scripts
- Thirty-three days later:
Agent = Model + Harness - “If you're not the model, you're the harness”
Loop engineering 迴圈工程
It never converged
- Neither of the two people who started it ever defined it
- Six accounts of five different objects — a person, a goal, a program, two machines, an organisation
- One vendor conceded in print that nobody agrees what a loop is
Graph engineering 圖結構工程
It was coined twice
- The careful, sourced essay sank. The twelve-word joke drew 3.1M views
- A survey had already named the practice in January: flow engineering
- The technical content is the least disputed of the five
How much of this is established every claim on this page, graded — sorted by standing, never mixed
| Standing | What is claimed | Where it comes from |
|---|---|---|
| Confirmed | The four coinage dates: 19 Jun 2025, 5 Feb 2026, 2–7 Jun 2026, 4 Jul 2026 | Each read at source; the two X timestamps re-derived from their post IDs for this page |
| Confirmed | The four movements: 102 days to invert, 33 to swell, 29 to fragment, 14 between the two coinages | Recomputed from the dated documents in this topic's research record, not copied from the articles |
| Confirmed | A January 2026 survey already called the graph practice flow engineering | arXiv:2601.12560 §5.1, read in full. The essay that coined “graph engineering” cites this paper for something else |
| Confirmed | Wikipedia has an article on the agent harness and none on context engineering | Both checked on 15 Aug 2026; the second returned HTTP 404. Loop and graph engineering were not checked |
| Confirmed, not verifiable here | Cherny's “my job is to write loops” — the sentence most often given as loop engineering's origin | Spoken at an event on 2 Jun 2026 and reaching this project only through a third party's write-up. Never heard at source |
| Confirmed, not verifiable here | Anthropic's two definitions — the augmented LLM, and “the smallest possible set of tokens” | anthropic.com refuses this session; the text was supplied by djTratoh and is summarised in input/summary/ |
| Unconfirmed | When “prompt engineering” was first said, and by whom | Never established. It is therefore excluded from the four-out-of-four pattern and from the 380-day span, rather than assumed to fit them |
| This project's reading | That each name arrives when a particular unit of work stops being the one a person controls | An inference drawn across the five records. No source states it, and it is marked as inference wherever it appears |
| This project's reading | That four gaps between practice and word amount to a pattern rather than a coincidence | Four instances with no counterexample found. Four is not many, and the page says so |
| Not public | The contents of the X Article that announced loop engineering's death | Behind X Premium. Title, byline and timestamp only — this page characterises none of its contents |
Timeline
How the vocabulary arrived the mechanisms are 2018–2022; the names are 2025–2026; there is very little in between
- 2018decaNLP casts ten tasks as questions over a context. The prompt is a task selector
- 2020GPT-3: 175B parameters, few-shot sometimes competitive with fine-tuning. The prompt now carries the examples
- Jan 2022Chain-of-thought: eight exemplars beat a fine-tuned model with a verifier. The prompt shapes the process — ten months before ChatGPT
- Oct 2022ReAct: reason, act, observe. The inner loop exists as a paper, four years before anyone names loop engineering
- 19 Dec 2024Anthropic publishes the augmented LLM — retrieval, tools, memory. This is context engineering, six months before it has a name
- 19 Jun 2025Context engineering is named, on X, as “providing all the context”. 182 days after the practice was published
- 25 Jun 2025Six days later a second definition: “just the right information for the next step”. Already narrower, and already per-turn
- 14 Jul 2025Chroma measures context rot across 18 models: they do not use their context uniformly, and get less reliable as it grows
- 29 Sep 2025“The smallest possible set of high-signal tokens.” The term has inverted in 102 days
- 26 Nov 2025Anthropic publishes the whole harness practice — an initializer, a feature file, a progress file. The word “harness” is used as an ordinary noun
- 18 Jan 2026A survey names the graph practice: it “is often described as flow engineering”. Nobody in the later material mentions this
- 23 Jan 2026OpenAI decomposes the inner loop in public, and calls its agent “our agent (or harness)” — the noun needing no defence, two weeks before it becomes a discipline
- 5 Feb 2026Harness engineering is coined, on a personal blog, for an instructions file and two scripts. Its author says he knows of no existing term
- 10 Mar 2026
Agent = Model + Harness. Thirty-three days, and the word now means everything that is not the model - 13 Apr 2026A paper formalises the loop as a scheduler whose ready set is exactly one — and calls its own proposal a Structured Graph Harness, three months before “graph engineering” is said
- 2 Jun 2026“I don't prompt Claude anymore. My job is to write loops.” A report of one person's practice, at an event. 130 days after the inner loop was published
- 7 Jun 2026An X post turns that report into an instruction — “you should be designing loops that prompt your agents” — and it travels. Neither post defines a loop
- Jun – 1 Jul 2026Six framings in 29 days, of five different objects: a person's habits, a goal, a program, two nested machines, a taxonomy, an organisation whose outer loop runs for weeks and contains users
- 4 Jul 2026Graph engineering is coined in earnest — seven minutes, six citations, three commitments. It sinks without trace
- 18 Jul 2026Twelve words on X, as a joke, at 00:34 UTC — 3.1M views. An obituary for loop engineering follows 4h35m later, behind a paywall. The joke is the version that travels
- 27–28 Jul 2026A Taiwanese trade title draws the five-layer ladder, dating every rung — and in the text beside it writes 目前還沒有定論, there is no settled conclusion
- 13 Aug 2026Wikipedia carries an entry for the agent harness, its etymology marked contested. There is still no article for context engineering, fourteen months after it was named
The Argument
The field's account of itself: a ladder asserted by the least hedged sources, and by nobody who did the naming
The standard account is a stack: prompt engineering inside context engineering, both inside the harness, the loop above that, the graph above the loop. It is a good story and it is drawn far more confidently than it is argued. The clearest statement of it is a diagram, which dates every layer and fixes the nesting; the article printed beside that diagram, by the same publisher on the same day, says in as many words that there is no settled conclusion about the order — some put the loop above the harness, and some leave the harness out altogether. One of those who leaves it out is the writer who coined graph engineering: his ladder has four rungs, and he omits the harness while leaning hardest on a paper whose proposal is called a Structured Graph Harness. The pattern is consistent enough to be worth naming. Confidence in this vocabulary rises with distance from the person who invented it.
- 5 vs 4rungs, depending on whether the harness is a layer or a container
- 4 / 4namers who hedged: “no existing term”, “I'm skeptical”, “should matter less over time”, and a joke
What the records show instead: a boundary moving outwards the reading is this project's; the movements underneath it are not
Read in date order, the five names do not stack. Each one arrives at the moment a particular unit of work stops being the unit a person can control. While the unit was one string, the discipline was prompt engineering. When it became one window, context engineering. When the window stopped being the edge of the system, everything outside the model — the harness. Then the repetition — the loop. Then what happens between windows — the graph. That reading is an inference and is marked as one; no source states it. What is not an inference is the consequence: five terms, and five different failure modes. Prompt engineering narrowed until it was a component. Context engineering inverted. Harness engineering swelled. Loop engineering fragmented into six accounts of five different objects. Graph engineering was coined twice, and the coinage that won was not meant seriously. Not one of the five simply came to mean one thing. A vocabulary that behaves this way is not a ladder being climbed. It is a field naming a boundary it keeps having to move.
- 182 → 71days between practice and word — the gap narrows as the pace picks up
- 41days apart, one person's two posts propelled two of the five names. Neither defines anything
What Others Add
What is actually agreed assembled across all five terms, and disputed by none of them
- 1Write at the right altitudeSpecific enough to guide, general enough to survive the next case. Not a list of every situation you have met
- 2Keep the tool set minimal and non-overlappingTwo tools that do nearly the same thing are a decision the model now has to get right, for no gain
- 3One job per nodeReached independently by three sources across two different terms. “A good node is boring”
- 4Verify from outsideThe system exists to disagree with the model's “I am done”. Asking a model to review its own answer is not external verification
- 5Checkpoint the state, not the transcriptState with a schema, written at every step. If memory is “whatever survived compaction”, failure at step 40 means starting again
- 6Put the budget in the stateThree sources, three terms, one sentence: if you cannot stop the system at a spend threshold, it is not autonomous — it is running up a bill
The evidence, and where it runs out the weakest part of the subject, and the field says so about itself
| Question | What exists | Standing |
|---|---|---|
| Does the harness matter? | One comparison with the right shape — model fixed, harness varied, on a third-party leaderboard. The same frontier model scores far apart in two harnesses | The strongest thing in the whole corpus, and it covers one term of five |
| Does a loop beat prompting? | Nothing. No controlled comparison was found in any source | An absence found by looking for it, not a gap in the search |
| Does a graph beat a loop? | Nothing. The one rigorous paper says of itself: “not a production implementation or empirical results”, under an explicit fairness disclaimer | Two terms running now rest on no measurement at all |
| Is the paper trail reliable? | Not yet. An encyclopaedia entry mis-dates its own primary source by a year; the one survey behind the graph case reports 42, 41, 49 and 68 in four places for a study stated as 70 | Both checked against the documents. Check dates and URLs yourself; this material is young |
Two facts about where this vocabulary comes from are worth more than any of the advocacy. The first: of the four coinages this project could date, none happened in a peer-reviewed venue — two X posts, a personal blog, a personal site and a conference talk between them. The papers are real and they arrive afterwards, adopting a word the field has already chosen. The second: the people who did the naming are the ones least sure about it. The man who coined harness engineering said he knew of no existing term and did not press the point. The man who named loop engineering wrote “its still early, I'm skeptical” in the same paragraph. The man whose definition of the harness everyone else restates argues that harnesses should matter less over time. And the post that made graph engineering famous was a joke, which a named analyst pointed out four days later — while insisting, correctly, that the joke worked because it pointed at something real. That last distinction is the one to keep. The naming is noise. The practice under it is not.
Conclusion
So what what to take from four names in 380 days
The practice comes first. Every time.
Four terms, four datable gaps, no counterexample: 182 days for context engineering, 71 for harness, 130 for loop, 167 for graph. In the graph case the practice was not merely being done — it had already been named, in a January survey, as flow engineering. If a new term appears, the useful question is not what it means but what has already been shipping under another name for the last six months.
This is a vocabulary, not a technology.
The mechanisms are 2018 to 2022 and they are published and peer-reviewed. The names are 2025 to 2026, and not one of the four datable coinages happened in a peer-reviewed venue. Almost nothing sits in between. That does not make the terms worthless — a field needs words — but it does mean a new name on this ladder is evidence about the conversation, not about the systems.
Trust the namers over the repeaters.
Every one of the four hedged — “I know of no existing term”, “I'm skeptical”, “harnesses should matter less over time”, and one outright joke. Downstream, the hedges vanish: the confident five-rung ladder is drawn by a publisher whose own article says there is no settled conclusion. Where a claim about this vocabulary sounds certain, check how far it is from whoever invented the word.
And the part that actually holds.
Right altitude, minimal tools, one job per node, external verification, checkpointed state, a budget you can stop the system at. Six practices, assembled from five separate arguments about what to call the layer they belong to, and contradicted by none of them. That is the durable part of this subject, and it would still be worth doing if every one of the five names were withdrawn tomorrow. It is also, so far, the part nobody has measured.