So how do we use “AI”?
I'm intending this to be the first in a series of posts about what I've learned over the last 6 months diving into the "AI Era", the technology behind it, and share some of my now-much-more-informed opinions around all of this. So I'm going to start with the big one I've asked previously...I think the real question here, in all of the chatter, is "How do we use this?". And here's what the current state seems to be: The people who make "AI" can't tell you how to use it. Not won't. Can't. The last few weeks seem to have made that plain to me.
Look at what the makers said when they had the world's attention yet again. The next generation of models will be "sobering for everybody" (OpenAI's CEO, September 3). "For too long the industry lied to people about the fact that this technology had risks" (Anthropic's CEO, on national television, September 13), with a call to slow the industry down. A trillion-dollar IPO shelved over safety. Chip stocks falling on the makers' own words. Notice the shape of every single statement: what the models will BE. Smarter, scarier, sooner. That is the half they can see.
Leading back to my original question: "How do we use this?" Not what it will become. What do we actually do with it, and trust it to do, now, on real work.
That's the question no lab can answer, because how to use these things is not a fact about the model. It's a fact about your organization, business, entity, etc: what right looks like there, what a mistake costs there, which corner the thing cuts on YOUR job, who checks its work before it ships. No lab can see any of that from where it sits. What they see is benchmarks and thumbs-up clicks, an ocean of feedback one inch deep. The answer gets built where the consequences live, or it doesn't get built at all.
One number still shows how much of it has been built: about 95% of enterprise AI pilots return nothing, and not one of those CEOs disputed it. The models got smarter all year. Has that number really moved? I think that number has nothing to do with the model either, it's our number. Because we (and I use that as a collective "we") don't really know how to use this. And guess what? That's ok! Because the people strapping a harness onto an LLM and calling it AI don't know either :P
So I'm dumping a series of posts, and articles, and all kind of stuff I've been working on as I've learned, experimented, built, destroyed, and delved into really fascinating academic topics. This is reality as I see it with AI right now: I have done, and can do, amazing things with it, but my ability to use it effectively in certain ways comes from a lot of experience in doing things those certain ways for a long time. But it does NOT work the way we have been sold, and most still believe, and we need to understand what we're dealing with before we can answer that most basic question: "How do we use this?"
So the themes of the subsequent posts. I'm trying to present them in plain English, with more detail in the background documents that go with each.
1. The model stops at the first answer it can confidently defend to you. Not when it's right.
2. AI doesn't need to be smarter. It needs a boss.
3. Rulebooks are theater. Boundaries and incentives are not.
4. Never let your AI grade its own homework.
5. Your corrections are the most valuable data in your company.
6. Rent the pipes. Own the judgment.
The labs will keep making the models smarter. The rest has to be built by the people doing the work.
The primer
- Primer 1 What is an LLM?
- Primer 2 What is a harness?
- Primer 3 What is context?
- Primer 4 What is “AI”?
The six themes
- Theme 1 The model stops at the first answer it can confidently defend to you. Not when it's right.
- Theme 2 AI doesn't need to be smarter. It needs a boss.
- Theme 3 Rulebooks are theater. Boundaries and incentives are not.
- Theme 4 Never let your AI grade its own homework.
- Theme 5 Your corrections are the most valuable data in your company.
- Theme 6 Rent the pipes. Own the judgment.
Sources
- "Sobering for everybody" (OpenAI's CEO, Axios, Sept 3)axios.com
- "For too long the industry lied" and the call to slow down (Anthropic's CEO, CBS Sunday Morning, Sept 13)cbsnews.com
- The IPO pushed to 2027 over safety (Fortune, Sept 12)fortune.com
- Chip stocks falling on those statements (Reuters, Sept 14)finance.yahoo.com
- The 95%: MIT NANDA, The GenAI Dividenanda.media.mit.edu