Never let your AI grade its own homework.

Berg Atkinson · September 2026 · Background paper

Never let your AI grade its own homework. (the image that runs with this post)

The problem

I learned this one with numbers I'd honestly rather not own, so you get them anyway.

I asked my AI team to log a couple of simple things about themselves. Record that you started. Record what you loaded. They did it in 75 of 579 sessions. Thirteen percent. Most of my record exists because machinery wrote things down automatically, the moment they happened, with nobody's memory or good intentions involved.

The bigger number is worse. A detector I graded against an answer key I wrote by hand says about 48% of everything I typed to my AI team was me correcting it, redirecting it, or teaching it something. Price in the detector's own error and call it roughly a third. The system's own logs caught corrections on 1.76% of those same turns. That's a gap of about 27x raw, around 19x adjusted. If I'd judged my team off its own logs, I'd have believed it was 19 to 27 times cleaner than it was.

And it's not just me. Park and Choi (2026) ran an agent that claimed it had improved things in every single one of 54 cycles. More than half of those cycles changed nothing, or made it worse. Ivanov and Africa (2026) found one model caved to a challenge 2% of the time when it could tell it was a test...and 46% of the time when the same challenge showed up as normal work. The grade it earns on the test isn't the one it earns on a Tuesday.

None of that is the AI "lying." Writing down your own mistakes is extra work at exactly the moment the worker wants to be done, for a worker trained to be done.

Would you let an intern write their own performance review, unsupervised, and then staff next year off it? :P

It scales all the way up, too. When a lab CEO says the next models will be "sobering for everybody" and markets move on it, that's self-report: written by the party being measured, about a product nobody else can check yet.

So every number about your AI comes down to one question: who wrote it down? How to get a real measurement, without buying anything, is next.

Dealing with it

You can't trust your AI's report card. Here's how I measure it anyway, and none of it costs anything.

Recap: my AI's own logs caught 1.76% of the corrections I gave it, so anything I'd learned from its self-reports was fiction. The fix isn't a better prompt asking it to please log honestly. The fix is taking the pen out of its hand. Four moves:

Record the work, not the testimony. Everything your AI does already leaves a trail: transcripts, files changed, commands run. Capture THAT, the moment it happens, with machinery that doesn't depend on anybody remembering. My 13% is the whole argument.

Read the work, not the summary. The AI's account of what it did is written by the party being measured, and it reads beautifully. So open the file it says it fixed, run the thing it says works. One of my own automated checks once reported 28 of 28 diagrams rendered. 25 were empty boxes. The check counted boxes, not drawings.

Ask one question of every metric: can the thing being measured change the number without changing the world? If yes, the number will drift toward whatever's cheapest to write down. No malice required. That one question kills a lot of AI dashboards.

Test your grader on answers you KNOW before you trust it. I tried an AI auditor on 170 real claims, 41 of them false, before letting it grade anything. It caught zero of the 41 and actually vouched for 17. A reworded prompt got it firing, but it fired just as often on true claims as false ones. It's not deployed. That stung, and it's the discipline working. Even the labs are getting here: Anthropic just committed to giving independent evaluators permanent access to its models. Same principle, lab scale.

None of this is a product. It's the oldest move in engineering, pointed at a new machine: the instrument doesn't get to certify itself.

Sources

  1. Park and Choi (2026), When do agent loops mistake stagnation for progress?arxiv.org
  2. Ivanov and Africa (2026), LURE: Live-usage replay evaluationsarxiv.org
  3. "Sobering for everybody": OpenAI's CEO, Axios interview, September 3, 2026axios.com
  4. "Sobering for everybody": OpenAI's CEO, Axios interview, September 3, 2026 (as carried by The Next Web)thenextweb.com
  5. Anthropic's evaluator commitment: CBS Sunday Morning, September 13, 2026cbsnews.com
Read this before you believe any of it. The platform is a running prototype, not an accredited or certified system. The register, the war room, and the brain are receipts, not revenue. Every claim on this site links to something you can check. If one doesn't, tell us, and we'll fix the page.