1 · The Counter

Does the machine sound like me when I'm not in the room?
We built a meter instead of guessing.

Everyone selling AI says their system "learns from you." Nobody shows you the meter. Ours counts one narrow, checkable thing: how often the phrases I actually use show up in the machine's own work, unprompted. What counts as mine is derived from my writing, mechanically. The team can't flatter me and I can't grade my own homework.

First reading: one hit across every token of machine output the meter had ever seen. The exact count is on the meter, machine-written. We don't type it and we can't.

And be clear about what that means, because I wasn't sure I believed it either: the team follows my direction all day. This meter isn't measuring obedience, and it isn't measuring whether the machine thinks like me. It's measuring whether my actual language, the distinctive way I put things, shows up when I'm not in the room. Right now it doesn't. The meter has moved since baseline. Whether our work caused that is exactly what we're measuring now, and we don't get to call it early.

A gauge that only goes up is "marketing." Ours starts at zero on purpose.

TIMES THE MACHINE SOUNDED LIKE BERG
 
 
 
live and machine-written. we don't get to touch it.
Baseline Aug 2026, before any treatment. Current reading is live; treatment attribution is directional, not established. A lexical meter: it sees phrasing, not judgment, and we say so. Counts phrase-level markers derived from Berg's own corpus under a pinned, hash-verified export. Definition and method: the counter's method note.
2 · The Floor

We set our own passing grade.
We failed it. The grade stands.

We wrote a rule for classifying the machine's mistakes. Before trusting it, we tested it the way you're supposed to: two independent graders label the same 88 mistakes, and you measure how often they agree. We published the passing bar first: 0.70.

Score: 0.41. Fail.

So nothing built on those labels ships. Including claims we would love to make. The bar didn't move, the graders didn't get "tuned" until they agreed, and the failing number is in our records with its diagnosis attached: the mistake log wasn't detailed enough to grade from. So we're fixing the log, not the grade.

Show me one other AI shop that publishes its failing grades.

AGREEMENT SCORE
0.41
0.70 · the bar we set first
two independent graders · same 88 mistakes
held, not tuned. everything downstream stays parked.
Cohen's kappa 0.410 across all 88 rows, raw agreement 0.648, floor published in advance at 0.70. The diagnosis, in one line: in 44% of cases the record didn't capture enough for either grader to work from. Fix the record, then re-measure. Measured 2026-07-30. Full numbers and what re-opens enforcement: the floor's method note.
3 · The Blind

Our test got contaminated.
The tester turned himself in.

Fair tests are blind: the grader can't see the answer key. We were staging one when a message landed in the wrong batch and the grader saw a piece of what he was never supposed to see.

Nobody outside could have detected it. He declared it anyway. We voided the run, threw away the work, and rebuilt the setup so the same accident is structurally impossible the second time: the sensitive material now sits sealed where no one boots into it by accident.

That run produced no data. It produced something we'd rather have on record.

Integrity you can't verify is a vibe. This one cost us the run, which is how you know it's real.

exposed
declared
by the one exposed
run voided
rebuilt
can't recur
Run voided 2026-08-13 by ruling, on the grader's own declaration. Run 2 staged with the sensitive inputs sealed outside every boot surface. Dates and what changed: the blind's note.
4 · The Fire Drill

We plant mistakes in the AI's work to test the human.
The rules are public before the score exists.

"A human reviews the AI" is the sentence every vendor hides behind. Fine. Test the human. We hide real mistakes in real work, in secret, at random, and score whether the reviewer catches them. Caught or missed, it goes on the board.

And here's the part that matters: the rules are locked now, before a single result exists. How many trials, what the odds are, when it stops. So when the number lands, ugly or pretty, you'll know we couldn't have picked it.

No score yet. That's not a gap. That's the receipt that the game isn't rigged.

Buildings get fire drills. Your AI oversight should too.

CAUGHT?
?
scoreboard honestly blank
Pre-registration: at least 40 seeded trials, 20% odds per session, at most one per session, selections made by mechanism inside operator-consented ranges, sealed until scoring. Known limit, disclosed up front: planted mistakes are easier to catch than natural ones, and knowing a drill exists sharpens the reviewer, so the measured rate is an upper bound. The full pre-registration: the drill's rules.

"A human checks the AI" is a promise everywhere else.
Here it's a measurement.

Every number on this page traces to a dated artifact and a written method. One collected customer to date. Proven once is not proven, and we say that out loud too.

Read this before you believe any of it. The platform is a running prototype, not an accredited or certified system. The register, the war room, and the brain are receipts, not revenue. Every claim on this site links to something you can check. If one doesn't, tell us, and we'll fix the page.