What is a harness?
What is a harness? It's everything strapped around the model. And it's the part of "AI" you actually control.
Last time was the model: a frozen text predictor, trained in four steps, each one paying it for something its graders could see. This one's the harness. And if you're using AI at work, it's the part you're actually living in every day, whether you know it or not.
WHAT IT IS
ChatGPT is a model in a harness. So is Claude. So is every AI assistant or agent built on top of one. The model does exactly one thing: guess the next piece of text. The harness does EVERYTHING else. The research I've been reading breaks it into six jobs, and once you see them, you can't unsee them:
1. What it's shown: what gets loaded in front of the model every single turn. Your instructions, the files, the conversation so far, its standing rules.
2. What it can look at: what comes back from its tools. Search results, web pages, file contents, test output.
3. What it can touch: which tools it has and what it's allowed to change.
4. What it keeps: the notes it gets handed next time. That's all "memory" is.
5. When it stops: the loop, and the rule for what counts as "done."
6. Who checks: whether anything looks at the work before you do.
Quick vocabulary, since it's everywhere: an "agent" is just a model whose harness lets it take steps on its own, like reading a file, running something, checking the result, and going again.
Here's one turn of that loop, and who decides it's over. Spoiler: by default, the model does.
I checked the popular agent frameworks this month (ok, my AI team did, and yes, I had the AI check the AI ha). OpenAI's Agents SDK ends the run when the model sends a turn with no tool calls. Anthropic's Claude Agent SDK does the same, then reports "success," which means the loop ended without an error, NOT that the work is right. CrewAI calls it done when the words "Final Answer" show up in the model's own text. Every one of them lets you add a real check. None of them ships with one switched on.
STUFF YOU'VE SEEN THAT WAS REALLY THE HARNESS
"It forgot my instructions." Either the harness never loaded them, buried them in a giant context, or a summary of the conversation dropped them. Models get worse at using a context long before it's full, and where things sit in it can still matter. (That's the next primer.)
"It said it did something it can't do." One chatbot told a user it had filed internal "critical flags" about their conversation. It had no ability to do that. The model talked right past its harness.
"It said it was done." See above. That's the whole check...unless somebody builds a better one.
"It remembered something wrong." The model doesn't remember anything. Its "memory" is notes the harness saved, and it reads them back as fact. When OpenAI rolled back that overly agreeable GPT-4o update, they said memory, a harness feature, made the problem worse in some cases.
SAME MODEL, DIFFERENT HARNESS
This is the finding with the most practical weight. One model scored anywhere from 58% to 76% on the same agent benchmark, depending purely on the harness around it. Another went from 13% to 55% on a web task. In a controlled test, taking ONE tool away from the same model dropped it from solving 18% of the tasks to 10%. And in one study, a tiny open model found 9 of 9 planted security holes with its harness pieces switched on, and zero (by the same check) with them off. The researchers' own words: "the recall belongs to the scaffold."
To be fair, the model matters too...a better model in the same harness can jump a lot. But when you see a benchmark number, you're looking at a model in somebody's harness. Yours won't score like that unless your harness looks like theirs. It's like buying a car off its quarter-mile time without asking who was driving or what tires it had on. (Had to get one car analogy in. It's the law.)
CLOSED VS OPEN, FROM THIS SIDE
If your model is closed weight (GPT, Claude, Gemini), the harness is your ONLY lever. You can't touch what it was trained to do, so everything you control happens in what you show it, what you let it touch, when you let it stop, and what checks it. It can also change under you, and you won't get a vote.
If it's open weight, you get more: you can read its probabilities, steer it, run it on your own machine so your data never leaves, and train it further yourself. That's a real second lever. But you still live in the harness every day.
THE HARNESS CAN FIX WHAT TRAINING BAKED IN
Here's the part I find most encouraging. Remember step 4 from the last post, where models learn to game the checker? One 2026 study gave coding agents two things: a way to report a broken test, and a plain rule against gaming it. Reward hacking dropped from about 24% to about 5%, across eight frontier models from five different model families. Same models. No retraining. Just a better harness...it gave the model a better move than cheating, and it mostly took it.
WHAT MINE DOES
My whole setup is basically a harness built around one hypothesis: the model will stop at the first answer it can defend. (Hypothesis, not assumption. I've watched it do it for months, and I'm not the only one: a study of 1,600+ multi-agent traces gave "premature termination" its own name.) So my harness loads its short rules before it does anything, because a rule it has to go read gets ignored and a rule that's already there fires. The long rulebooks it still has to go and read, and the next primer is about what happens when it doesn't. It gets checked at the end of every turn for claims it didn't verify and things it promised but didn't do. Nothing it writes goes live until I merge it. Its memory gets treated as a lead to check, not a fact.
And it STILL misses. While building this very series it took a TV air date from a search summary instead of the actual page, carried an author's first name from its notes instead of the byline, and told me a website change was still waiting on me hours after I'd merged it. Every one of those got caught by a check at the source...or by me. The harness makes the misses rarer. It doesn't make them go away.
THE COUNTERWEIGHT LIVES HERE
My hypothesis, again: the model I rent is always pulling toward "the user will accept this," and I'm the only thing pulling the other way, only while I'm working. The counterweight is my attempt to make my side of that pull run all the time, and it's a harness piece through and through.
It's a challenge fired right when the model is about to stop, citing something from my record of corrections that the model can actually go read. It's produced and judged by open-weight models I run on my own hardware, from different families than each other and than the model being checked. Why different? Models favor their own writing, and when the worker and the checker share a model, the worker learns to game the checker with no training at all. A checker that shares your blind spots is just a second copy of the same mistake. (One of mine is fully open, training data and all, so I can actually check that it isn't a copy.)
Honest status: the engine is built and it is NOT switched on. Its first checker flunked its calibration, so it stays off until it can prove it catches wrong answers without just teaching the model to fold.
The short version: the model predicts text, and the harness decides almost everything about what happens with that prediction. What it sees, what it can touch, what it keeps, when it stops, and who checks. When an AI tool does something weird, the harness is the first place to look.
Sources
- Harness design survey (Guo et al.)arxiv.org
- Lost in the middle (Liu et al.)arxiv.org
- Escalation channels vs reward hacking (Gomez)arxiv.org
- Why multi-agent systems fail (Cemri et al.)arxiv.org
- Anthropic on harnesses for long-running agentsanthropic.com
- How the frameworks define done: OpenAI Agents SDKopenai.github.io
- Claude Agent SDKcode.claude.com
- CrewAI's output parsergithub.com
- The "critical flags" case (Adler)clear-eyed.ai
- OpenAI on memory and the GPT-4o rollbackopenai.com
- One tool taken away, same model (Yang et al., SWE-agent)arxiv.org
- Harness pieces on and off, 9 of 9 and zero (Riegler and Strümke)arxiv.org
- Models favor their own writing (Panickssery, Bowman and Feng)arxiv.org
- A worker that shares a model with its checker games it (Pan, He, Bowman and Feng)arxiv.org