Rulebooks are theater. Boundaries and incentives are not.
The problem
Every AI rollout gets here by week three: the rules document. Usage policy, style guide, "the AI must always," "the AI must never." I did it too, ha. My setup carries well over 150 written rules for my AI team, and I'd LOVE to tell you the adherence rate. I can't. I've never measured it, so I'm not going to make one up.
What I can tell you is what my own logs show. One day in July, seven of those rules got broken. Seven misses, one day. The page my AI reads at startup says it better than I can: "A rule asks a future session to remember." That's the whole problem. A rule the AI has to go read competes with the actual work for its attention, and it loses.
It gets better (worse?). I track two kinds of mistakes closely: blaming the wrong cause, and claiming it checked things it hadn't. They came back in the very next session 48 times out of 75. Once, it read the note describing that exact mistake from a few days earlier, acknowledged it...and repeated it within minutes. The lesson was written down. It just wasn't anywhere that mattered.
Before you blame the model, though, the research says this is the STABLE outcome. Wang and Huang (2026) proved that any optimized agent will under-invest in whatever you don't actually evaluate. Not a bug. An equilibrium. Everything you check pulls it one way, and your document doesn't pull at all.
You don't beat an equilibrium with a memo.
So if writing rules down doesn't govern the thing, what does? That's the next post, and I learned those answers the hard way.
Dealing with it
Rules don't govern your AI. Here's what does, and it's cheaper than the rulebook was.
Recap: rulebooks are theater for a worker that's born new every morning, and drifting away from them is the stable outcome, not a mood. Three things actually hold, all of them showing up in my own shop.
First, WHERE the rule sits. Same words, delivered differently, and it fires. A rule sitting in a document has to win a fight for the model's attention against the work, and it loses. A rule loaded into its context before it starts doesn't have to win anything. It's just there, like a job description. Research points the same way: re-injecting the goal mid-conversation cut drift by 7 to 12% across three models (Dongre et al., 2025).
Second, something that can say NO. A rule says please don't. A boundary refuses: a check that blocks the bad move at the moment it's attempted, and writes down that it did. Validation before anything sends, checks before anything merges, hard stops on numbers I've retired.
Third, incentives, meaning what actually gets checked. The equilibrium from the last post has a flip side: bring something under evaluation and the behavior moves. A memo changes nothing that gets checked. A boundary does.
Now the honest part, because I promised receipts. I joined my own boundary checks to the corrections I actually made over one window in September. They ran 674 times and refused 93 times...and only 2 of those refusals lined up with a real mistake I went on to correct. 14 of my corrections landed while a check was running and said nothing. So yes, boundaries hold where rules drift, but a boundary aimed at the wrong thing is just noise with a badge. Aim it at the mistakes you actually get.
So audit your own setup this week. For each rule you care about: does it live where it LOADS, or just where it's stored? If the AI breaks it anyway, can anything say no? And is that "no" aimed at the mistakes you really see?
Spend the memo budget on the humans. They're the only ones in the building who'll read it twice :)