Blog
Plain-English writing on using AI for real work. Every post links the paper behind it, with every source.
The primer
-
Primer 1What is an LLM?
What is an LLM? A text predictor. That's it. Trained in four steps, and every one of those steps paid it for something...which is exactly where the habits you fight with come from.
-
Primer 2What is a harness?
What is a harness? It's everything strapped around the model. And it's the part of "AI" you actually control. Last time was the model: a frozen text predictor, trained in four steps, each one paying it for something its graders could see. This one's the harness.
-
Primer 3What is context?
When a conversation with an AI gets too long for the model, one of a few things happens depending on the tool: the early part quietly stops counting, or it tells you to start a new chat, or a model writes a summary and the work carries on from it (that last one is called compaction).
-
Primer 4What is “AI”?
So what IS "AI", when people say it right now? The first two primers took it apart: first the model, then the harness. Put them back together and you have what almost everyone means by "AI" right now: two things wearing one name.
How do we use this?
-
The lead-inSo how do we use “AI”?
I'm intending this to be the first in a series of posts about what I've learned over the last 6 months diving into the "AI Era", the technology behind it, and share some of my now-much-more-informed opinions around all of this.
-
Theme 1The model stops at the first answer it can confidently defend to you. Not when it's right.
This is the one that made everything else click for me, so it goes first. Once you see it, a lot of the "weird AI behavior" you deal with every day stops being weird.
-
Theme 2AI doesn't need to be smarter. It needs a boss.
About 95% of enterprise AI pilots return nothing, per MIT's NANDA project, after 30 to 40 billion dollars of enterprise investment. And here's the part I keep coming back to: the report itself says that divide is "not...driven by model quality or regulation." It's "determined by approach."
-
Theme 3Rulebooks are theater. Boundaries and incentives are not.
Every AI rollout gets here by week three: the rules document. Usage policy, style guide, "the AI must always," "the AI must never." I did it too, ha. My setup carries well over 150 written rules for my AI team, and I'd LOVE to tell you the adherence rate. I can't.
-
Theme 4Never let your AI grade its own homework.
I learned this one with numbers I'd honestly rather not own, so you get them anyway. I asked my AI team to log a couple of simple things about themselves. Record that you started. Record what you loaded. They did it in 75 of 579 sessions. Thirteen percent.
-
Theme 5Your corrections are the most valuable data in your company.
They're free, they're already being written...and odds are you're deleting them. Every time you correct an AI, you're writing down a labeled example of your judgment on your work: what was wrong, why, and what right looked like. No lab can buy that.
-
Theme 6Rent the pipes. Own the judgment.
The AI plumbing is going commodity, right on schedule. Managed agent platforms are here: the orchestration, the retries, the state handling, all the stuff everyone dreaded building now rents by the month. Good!