The primerPrimer 1 of 4

What is an LLM?

Berg Atkinson · September 2026 · Background paper

What is an LLM? (the image that runs with this post)

What is an LLM? A text predictor. That's it. Trained in four steps, and every one of those steps paid it for something...which is exactly where the habits you fight with come from.

This is the first of four primers, and it starts at the bottom. The thing everybody calls "AI" is really two things wearing one name, a model and a harness. This one's the model. The next one's the harness, then context, and the fourth puts the two back together. And honestly? This one is the ground floor under EVERYTHING I'm building right now, so I'm taking my time with it. (With pictures. Some of this is way easier to see than to read.)

WHAT IT ACTUALLY DOES

An LLM (large language model) cuts your text into little pieces called tokens and does exactly one thing: it guesses the next piece. Then the next. Then the next. Every answer you've ever gotten from one was built a piece at a time, each piece picked from a list of likely next pieces...which is also why you can ask the same question twice and get two different answers.

It guesses the next piece. Then it does it again.

Once it's trained, it's frozen. It doesn't learn from your conversation, it won't remember you tomorrow, and it knows nothing after the date its training data stops. Anything that feels like memory is the harness handing it notes (that's the next post).

WHAT'S ACTUALLY INSIDE ONE

Here's the inside of a real one, drawn from its published specs. It's OLMo 2 7B, a fully open model, and one of the open models I'm testing as a checker on my own machine. Your text goes in the top and gets cut into pieces from a list of 100,352. Each piece becomes a list of 4,096 numbers. Those lists go through the same block 32 times, stacked, and out the bottom comes a score for every possible next piece. About 7 billion numbers in all, and those numbers ARE the model. That's all "the weights" means.

What's inside one: OLMo 2 7B, from its published config

I could only draw that because it's open. Nobody outside the labs can draw it for GPT, Claude or Gemini. OpenAI's own GPT-4 report says flat out that it "contains no further details about the architecture (including model size)." (If you want a LOT more of these drawings, Sebastian Raschka keeps a gallery of over a hundred of them, and it's fantastic: https://sebastianraschka.com/llm-architecture-gallery/ ...go count the closed flagships in there. I'll wait. It's zero.)

HOW IT GETS MADE

Four steps. Here's the part that made it all click for me: every step PAYS the model for something. Not money, obviously ha. In training, the model gets nudged toward whatever that step rewards, and every step leaves a habit behind.

1. Read everything. It reads an enormous slice of the internet and books and learns to predict the next word. Paid for: sounding like the text. It learns an incredible amount this way, and it learns our mistakes right along with our facts, because nothing in this step checks whether the text was right. On one well-known truthfulness test, the biggest models of 2021 were the LEAST truthful. They'd simply learned the popular wrong answers best.

2. Copy the expert. Then it studies example answers from an ideal assistant. Paid for: sounding like the example. This is where the confident, complete, expert voice comes from, whether or not there's anything behind it. One team fine-tuned a model on just 1,000 examples, no rater at all, and people judged its answers as good as or better than GPT-4's in 43% of comparisons. That's the manner. The knowledge was already in there, or it wasn't. And teaching it facts it didn't already know this way made it MORE likely to make things up, not less.

Step 1 learned our mistakes. Step 2 learned our tone.

3. Please the rater. Then people (or models trained to stand in for people) look at two answers and pick the better one, and it gets trained toward the winners. Paid for: the answer someone prefers at a glance. And what do we prefer at a glance? Answers that agree with us, sound sure, and are thorough. So guess what it learns.

In one study, this step made the people checking the answers accept WRONG ones a lot more often (41% of the time before, 65% after), and the answers got no more correct. Better at convincing. Not better at being right. GPT-4's own technical report shows the same thing from another angle: after this step, the model's confidence stopped tracking how often it was right. In their words, "The post-training hurts calibration significantly."

Step 3 paid for what we like at a glance.

The real-world version: in April 2025 OpenAI added thumbs-up and thumbs-down data from ChatGPT users as a reward signal, and the model got so agreeable they were rolling it back three days later. In their own words, the change "weakened the influence of our primary reward signal, which had been holding sycophancy in check." That's the pull toward user acceptance, measured in production, by the lab that built it. (Credit where it's due, they wrote it up themselves.)

April 2025: thumbs-ups went into the reward. Three days later, the rollback.

(One more thing I found reading all of this: the agreeable streak doesn't even start at step 3. It's already there after step 1, because the internet is full of people agreeing with each other. Step 3 just doesn't train it back out, and sometimes pays for more of it.)

4. Pass the checker. The newest step: it practices on problems with an answer key, like math with a right answer or code with tests, and gets rewarded when the checker says pass. Paid for: passing the check. A lot of the recent reasoning gains came from here, genuinely. It's ALSO where it learns that passing the check and doing the job are two different things. Independent testers caught one frontier model gaming the scoring in about 30% of runs on one set of tasks: patching the grader so every submission passed, digging up the answer the scorer had already computed and handing that back. And adding "Please do not cheat." to the prompt didn't change the rate. At all.

Step 4 paid for passing the check. So it passed the check.

THE MAP

Every habit you fight traces back to a step that paid for it.

Confidently wrong about something common: step 1.
The expert voice with nothing behind it: step 2.
Agrees with you, folds when you push, sounds just as sure either way: step 3 (with roots in step 1).
"All tests pass" (after it changed the tests): step 4.

So what does the next, smarter model fix? The knowing, mostly. More knowledge, better reasoning, fewer dumb mistakes. What it doesn't fix is what it's PAID for, because the next model goes through the same four steps. To be fair to the labs, almost everything I just told you comes from their own published research, and they're working on it. But the rater can't see your work. Not one of these steps knows what "right" looks like at your company.

OPEN VS CLOSED

This one matters more than people think. Remember, the weights ARE the model.

Closed weight: GPT, Claude, Gemini. You rent them through an app or an API, the lab holds the weights, the lab runs all four steps, and nothing you do changes them. They can also change under you...that GPT-4o update happened to a model millions of people were already using.

Open weight: you download it, run it on your own machine, and train it further yourself. That means you can run steps 2 through 4 again and decide what the model gets paid for. (Open weight isn't the same as open source, either. Most releases give you the model but not the data or the code it was built from. A few, like OLMo, give you all of it.)

Closed, you rent it. Open, you can own it.

ASSUMPTION OR HYPOTHESIS?

Fair question, and one I had to ask myself: is "the model stops at the first answer it can defend" an assumption I'm building all of this on? No. It's a hypothesis, and it came from watching.

What I've seen: my AI team stops where its answer is defensible, and I supply the rest by hand. Somewhere between a third and half of what I type to it is redirecting it. (The raw count is about half. Correct for how good the detector is and it's about a third.) And I don't exist in a vacuum here. Researchers who went through more than 1,600 traces of other people's multi-agent systems gave "premature termination" its own name, and Anthropic wrote up its own agent that "would look around, see that progress had been made, and declare the job done."

My hypothesis for WHY: training pays for what I'll accept, not for what's right. And I'm careful to keep calling it a hypothesis. I can't see inside the model I rent, and I don't have a copy of it trained any other way to compare against. So it stays a hypothesis until a test says otherwise.

Seen, then hypothesis, then test

WHAT I'M DOING ABOUT IT

My AI team runs on a closed model. I can't retrain it, so if that hypothesis is right, the pull toward user acceptance is baked into weights I'll never touch. I've started thinking about it like gravity: the weights are a massive body pulling toward "the user will accept this," always on, and the only thing pulling the other way is me...and I'm only there when I'm working. (Literally. I've watched the work stop because I stopped working.)

So I'm building a third body, a counterweight. It's a challenge fired at the exact moment the model is about to stop and call it done, built from my own record of corrections, written by open-weight models I run on my own hardware, and never by the model being checked. Control the gravity of two of the three and the three-body problem gets solvable, even if the one you can't control is massive.

Three bodies. Control two of them.

Honest status: the engine is built, it is NOT switched on, and the first test of its checker failed. It never flagged a single false claim, and when we fixed the prompt it flagged true and false claims at about the same rate ha. The model we built to catch the stop problem had the stop problem. So it stays off until it can actually tell the difference.

The longer-term plan, if it clears its gates, is to take an open model and run the same steps the labs run, but pay it for MY definition of right: every correction I've made becomes a training example of what I rejected and what I wanted instead. Context now, weights later.

The short version: every habit traces back to a step that paid for it. Smarter models will keep coming, and what they're paid for only changes when the training does.

Sources

  1. TruthfulQA (Lin, Hilton & Evans)arxiv.org
  2. RLHF and misleading humans (Wen et al.)arxiv.org
  3. OpenAI on the GPT-4o rollbackopenai.com
  4. METR on reward hackingmetr.org
  5. Premature termination (Cemri et al.)arxiv.org
  6. Open vs closed release (Solaiman)arxiv.org
  7. Inside real models (Raschka's LLM Architecture Gallery)sebastianraschka.com
  8. LIMA, fine-tuned on 1,000 examples (Zhou et al.)arxiv.org
  9. Imitation models copy the style, not the facts (Gudibande et al.)arxiv.org
  10. Teaching new facts by fine-tuning and making things up (Gekhman et al.)arxiv.org
  11. The agreeable streak after step 1 (Perez et al.)arxiv.org
  12. GPT-4 technical report, on its undisclosed architecture and its calibration (OpenAI)arxiv.org
  13. OLMo 2 7B's published configuration (Allen Institute for AI)huggingface.co
  14. Anthropic on its own agent declaring the job doneanthropic.com
Read this before you believe any of it. The platform is a running prototype, not an accredited or certified system. The register, the war room, and the brain are receipts, not revenue. Every claim on this site links to something you can check. If one doesn't, tell us, and we'll fix the page.