Does the machine sound like me when I'm not in the room?
We built a meter instead of guessing.
Everyone selling AI says their system "learns from you." Nobody shows you the meter. Ours counts one narrow, checkable thing: how often the phrases I actually use show up in the machine's own work, unprompted. What counts as mine is derived from my writing, mechanically. The team can't flatter me and I can't grade my own homework.
First reading: one hit across every token of machine output the meter had ever seen. The exact count is on the meter, machine-written. We don't type it and we can't.
And be clear about what that means, because I wasn't sure I believed it either: the team follows my direction all day. This meter isn't measuring obedience, and it isn't measuring whether the machine thinks like me. It's measuring whether my actual language, the distinctive way I put things, shows up when I'm not in the room. Right now it doesn't. The meter has moved since baseline. Whether our work caused that is exactly what we're measuring now, and we don't get to call it early.
A gauge that only goes up is "marketing." Ours starts at zero on purpose.
We set our own passing grade.
We failed it. The grade stands.
We wrote a rule for classifying the machine's mistakes. Before trusting it, we tested it the way you're supposed to: two independent graders label the same 88 mistakes, and you measure how often they agree. We published the passing bar first: 0.70.
Score: 0.41. Fail.
So nothing built on those labels ships. Including claims we would love to make. The bar didn't move, the graders didn't get "tuned" until they agreed, and the failing number is in our records with its diagnosis attached: the mistake log wasn't detailed enough to grade from. So we're fixing the log, not the grade.
Show me one other AI shop that publishes its failing grades.
Our test got contaminated.
The tester turned himself in.
Fair tests are blind: the grader can't see the answer key. We were staging one when a message landed in the wrong batch and the grader saw a piece of what he was never supposed to see.
Nobody outside could have detected it. He declared it anyway. We voided the run, threw away the work, and rebuilt the setup so the same accident is structurally impossible the second time: the sensitive material now sits sealed where no one boots into it by accident.
That run produced no data. It produced something we'd rather have on record.
Integrity you can't verify is a vibe. This one cost us the run, which is how you know it's real.
We plant mistakes in the AI's work to test the human.
The rules are public before the score exists.
"A human reviews the AI" is the sentence every vendor hides behind. Fine. Test the human. We hide real mistakes in real work, in secret, at random, and score whether the reviewer catches them. Caught or missed, it goes on the board.
And here's the part that matters: the rules are locked now, before a single result exists. How many trials, what the odds are, when it stops. So when the number lands, ugly or pretty, you'll know we couldn't have picked it.
No score yet. That's not a gap. That's the receipt that the game isn't rigged.
Buildings get fire drills. Your AI oversight should too.
"A human checks the AI" is a promise everywhere else.
Here it's a measurement.
Every number on this page traces to a dated artifact and a written method. One collected customer to date. Proven once is not proven, and we say that out loud too.