Reading view

There are new articles available, click to refresh the page.

Etzioni on AI: What kids tell chatbots, but not you

Teens are taking questions about their health, their moods and their friendships to chatbots — often without telling anyone. (GPT-5.6 Sol)

Kids are heading back to school across the country with a tool in their pocket that will do their homework, answer the questions they’re too embarrassed to ask an adult, and never tell anyone they asked.

MIT Technology Review asked kids aged 10 to 18 what they make of AI. Their answers had me plowing through every survey and interview with kids about AI I could find. Plenty of what kids do with it is ordinary: looking things up, homework, messing around.

This article is about the parts they keep to themselves.

1. Kids are taking their mental health to a chatbot, and telling nobody.

Nearly one in five Americans aged 12 to 21 has asked a chatbot for help when feeling sad, angry or nervous. RAND puts it at 8.2 million young people, and 63% of them told no one at all.

The silence is a calculation about what adults would do, as one teenager explained to YouGov’s researchers.

The most vulnerable kids use chatbots more.

  • Internet Matters found that British children with a health condition requiring professional help are nearly three times as likely to use companion bots like Character.AI or Replika.
  • A third of all child users say chatting with a bot feels like talking to a friend, and among the vulnerable that figure goes up to half.
  • One in eight of all child users says they talk to a chatbot because there’s no one else, and among the vulnerable it’s nearly one in four.

2. Kids see both sides of AI in school.

Some teenagers see AI as a shortcut past the hard parts of schoolwork.

Others describe using it the opposite way.

Between May and December 2025, RAND watched homework use among middle school, high school and college students climb from 48% to 62%. Over a slightly longer stretch, from February, the share saying AI harms critical thinking climbed from 54% to 67%. The same students can hold both thoughts at once.

The Concord Monitor interviewed eight New Hampshire high schoolers this spring. They weren’t outraged at their classmates, but they described losing motivation to do the work themselves. One in ten teens tells Pew they do all or most of their schoolwork with a chatbot’s help.

Caledonia Mahon, a Concord High senior, watched a classmate stand up and admit he’d had ChatGPT write a personal reflection, and started wondering why she was still writing her own.

Her classmate Andrew Pfitzenmayer supplied the detail I can’t get out of my head, about a boy in his advanced history class.

Yet, the same technology can be positive for kids who don’t have much else.

By May 2025, 84% of American high schoolers were using AI for schoolwork at least occasionally. The kids who lean on it hardest have the least support around them, which is why telling them to quit doesn’t work.

If quitting isn’t the option, what’s left is changing what the tool does when a kid opens it. Two Seattle startups are trying that.

  • Wild Zebra, started in 2024 by Edan Shahar and Erik Selberg, makes a math and reading tutor for grades 2 through 9 that answers a stuck student with a question instead of a solution. It’s gone from about 6,000 students in four pilot schools a year ago to tens of thousands now, and raised $6 million in August.
  • Maximal Learning, built by Microsoft veterans Eran Megiddo and Liviu Asnash, ships an app called Wick that coaches planning, time management and study habits rather than producing homework. Neither company pretends teenagers will stop bringing AI to their assignments. Both are built so the student still has to do the thinking.

3. Kids are writing the rules adults haven’t.

Suspicion of AI runs strongest among the kids furthest from it: 78% of students who don’t use AI say it harms critical thinking, against 60% of those who do.

In July, 98 high schoolers from all 50 states spent three days in a replica of the U.S. Senate chamber in Boston and passed what Congress hasn’t. Their STUDENTS FIRST Act requires AI literacy, bars teachers from letting AI decide a grade on its own, lets a teacher ask a student to defend suspected work out loud, and gives any student the right to an alternative to an AI assignment.

A majority of educators told Education Week’s Research Center last fall that their district has no AI policy, or that they don’t know whether it does. At least seven states have a comprehensive one, by Education Week’s count.

Ashley Kannan has taught social studies for almost 30 years in Oak Park, Ill. He planned to ignore AI, then noticed that every conversation about it in his building was about catching students. He recruited nine eighth graders to spend a year working out what AI was actually good for. In April they briefed 60+ teachers and parents on the rules they thought were needed.

Hazel is 17, a rock climber in New York who wants to be an ecologist. She won’t touch AI, and her reason is specific: what data centers do to the water supplies of the towns they’re built in. Asked what advice she’d give other kids, she gave four words.

Hazel can afford that advice. Most kids can’t.

So the practical alternative is narrower than just telling them to quit: ask them what they use AI for, ask what they’ve decided not to use it for, and don’t make an honest answer cost them the tool. Right now 61% of young people say their parents rarely or never talk to them about AI, and 53% say the same of their teachers.

Hazel may walk away from AI. But Amira, describing a pain in her side to a chatbot, may not, and the odds are nobody in her house knows she’s doing it.

Bill Gates in his own words: How he’s using AI, and why he’s worried about the future

Bill Gates, shown here in April 2025, released a memo this week warning that the world isn’t ready for AI. (GeekWire Photo / Kevin Lisota)

This week on the GeekWire Podcast: Bill Gates published a new essay warning that the AI industry is crossing the safety lines it set for itself, and that nobody is preparing for what’s coming. At age 70, he also uses AI more than most people half his age, and he finds it enthralling, as you’ll hear on this week’s show, with highlights from our interview with him.

Along the way, we dig into his three proposals: new institutions for managing the transition, a category of jobs reserved for humans, and a tax on the use and purchase of AI and robots.

The change in his own tech usage: “I joke with people that I used to have Claude-like people that I would send email to, but they were so slow, and there were some topics they didn’t actually know. … It’s three a.m. I want to understand sodium batteries, and now there’s no reason to go to sleep. Here we go. Yeah, it’s crazy.”

How he uses AI specifically: “If you’re a curious person, this is a mind-blowing time. When I’m working on malaria, nutrition, my poor humans that I work with always get these long conversations from me, where I paste in — me, Claude, me, ChatGPT. Sometimes I do it if there’s three of us: Claude, ChatGPT and me, debating these things.”

On where personal agents are headed: “We will get to a point where you won’t buy things yourself. You just won’t. … You won’t go to those applications. You’ll just go to your personal agent. … From a productivity point of view, we are in heaven.”

What has surprised him: “I was shocked by ChatGPT, and I was shocked by Claude Code. Those are both things where I went, oh my God. … I did not expect that a statistical machine would essentially learn to read, and the idea that the code is better than human code. Those are two stunning thresholds.”

On writing this essay: “It’s very unnatural for me to think that innovation may be a net negative if it’s not managed properly. The more I wrote the memo, the more I was like, Jesus, we really need to get our act together here. Even though this may come across as negative, that’s the truth. If we don’t step up, the negatives will substantially outweigh the positives.”

What AI leaders say privately: “You’re in this perverse period right now where people in the AI industry who are willing to say that AI might have some negative effects are told, ‘Hey, you’re hurting our PR while we’re trying to raise trillions of dollars.’ … I know they’re all worried. Or all of them that I know, which is basically everybody but Elon.”

On losing control of AI: “The wake-up for the memo is that the bad stuff thresholds are all being crossed. Even lack of control that I thought would be many years from now, we’re seeing lack of control. … These are people who are super expert on the thing, going, well, maybe we won’t be able to control these things. What kind of risk have we chosen to run here?”

On how fast robots are coming: “What’s weird about AI is it’s better at doing jobs across the entire economy, including physical jobs when the robots come — which you can guess when that is, but my view is it’s only a couple of years.”

Is he still an optimist? “I don’t think being pessimistic is helpful. I do think, wow, this is sure an interesting time. I’m the guy who in my 30s thought people in their 50s or 60s didn’t understand anything. So it’s kind of bizarre if a guy who’s 70 comes and writes a memo that’s actually helpful. … But I am very concerned. And honestly, when you get people one-on-one, so are they.”

Related headlines and links

Subscribe to GeekWire in Apple Podcasts, Spotify, or wherever you listen.

Edited and produced by Curt Milton. Music by Daniel L.K. Caldwell.

Etzioni on AI: Bill Gates has the right diagnosis but the wrong prescription

Bill Gates, whose new essay warns of the risks ahead in the AI era, during a 2017 interview. (GeekWire File Photo / Kevin Lisota)

When Bill Gates talks, people listen. This week he published a lengthy essay on what AI is going to do to work, and told GeekWire that people inside AI companies who name the downsides get told, “Hey, you’re hurting our PR while we’re trying to raise trillions of dollars.”

He’s right about the hard part. The job displacement he describes lands on young workers first, and the safety net is funded by taxes on the very wages that AI erodes. He prescribes three treatments: new institutions at home and abroad, a tax on AI tokens and robots, and “Human Reserved,” a category of jobs only people may hold.

Gates has the diagnosis right but the prescription mostly wrong. I’d sign the robot tax tomorrow, because hiring a person costs you payroll tax every year while buying a robot gets written off in year one. The other two I’d send back.

Let’s start with what’s solid. Stanford’s Digital Economy Lab updated its “Canaries in the Coal Mine” work this month. Employment for 22-to-25-year-olds in the most AI-exposed occupations is running 19% below where it would be if it had kept pace with their peers in less exposed work, up from 15% a year ago. The same authors say they don’t see widespread, economy-wide displacement, and unemployment held at 4.1% in July.

The AI damage isn’t arriving as layoffs. It’s arriving as jobs that never get posted, and Gates is right that the young get it first.

Now the token tax. Tokens (essentially words) are what AI companies bill by. Taxing tokens is like taxing keystrokes: it measures effort, not displacement.

A high school class working through calculus with an AI tutor burns tokens continuously. A model that quietly retires a 40-person customer center might burn relatively few. The tax lands hardest on the uses Gates says he wants to protect.

Stanford’s AI Index put the cost of GPT-3.5-level performance at $20 per million tokens in November 2022 and seven cents by October 2024, a 280-fold drop. You’d be indexing the safety net to a number that falls every year while displacement rises.

And you can’t collect it. Inference runs on laptops and phones now, and on servers in whatever country declines to sign. A token tax is a tax on whoever uses an American API, and every dollar it adds makes a Chinese model look cheaper. We’d be slowing ourselves down and not China.

Gates says the institutions will take years to build, and also says we can’t afford to move slowly. He’s right twice, and that’s the problem. He wants the international body to borrow from nuclear inspections and aviation regulation. That may pan out in the long term, though the UN is the cautionary tale for the bureaucratic nightmare that the international community can produce.

Meanwhile we have functional agencies with jurisdiction today. The FDA can rule on AI in diagnosis. The FTC can go after AI-enabled fraud. We don’t need a new agency to say a bank can’t deny your mortgage because a model felt like it. We need the banking regulator to reiterate it forcefully.

That leaves Human Reserved, his best idea but his most privileged one. Gates would protect a job for either of two reasons: the role is deeply personal, like a caregiver, or the people who hold it are unlikely to find other work. Only one of those holds.

Freezing headcount because the workers have nowhere else to go protects the job for a while and makes the service more expensive along the way. Reserving the moments when a human being is the point is defensible, and Gates makes that case well. On a robot delivering the news that you have an incurable disease, he writes, “There’s no technical reason why it couldn’t,” and adds, “Yet it shouldn’t.” He’s right.

I made the case in WIRED nine years ago that displaced workers should move into caregiving, and that it would take real money to lift the pay enough to draw them.

The problem with Human Reserved is that it assumes there’s a human being available. Home health and personal care aides earn a median of $34,900 a year, and BLS projects roughly 765,000 openings in that occupation every year through 2034. At that wage, they keep coming open. A third of home care aides are immigrants, and tighter enforcement threatens that supply. A rule that reserves care for people, in a market with no spare people, reserves care for the families who can outbid everyone else.

Gates half-anticipates this, telling The New York Times he might be a flawed messenger because of his wealth. On this point he is. The caregivers who gave his father something irreplaceable were in that room because someone could pay them to be there.

So don’t fence AI out of the room. Put it to work in the hours nobody is paid to cover.

In February the Times ran Eli Saslow’s story about Jan Worrell, 85, living alone on Washington’s Long Beach Peninsula with an AI companion called ElliQ that engages her about eight times a day and pushes her to stay hydrated and moving. (I serve on ElliQ’s board, and I joined because the company builds a machine that extends a caregiver’s reach instead of replacing one.)

Her goal, she told her doctor, was to never live anywhere else. Fund enough aides to cover the hours that need a person and put the machine on the rest.

Here’s where I net out: equalize the tax treatment of labor and capital, which Congress could do next session, and route the proceeds into retraining and into topping up the pay of workers who land in lower-paying jobs. That’s a better answer than a protected job title.

Drop the token tax, build the caregiving workforce instead of fencing it off, and use the regulators we already have while somebody works on the ones we don’t.

Etzioni on AI: An Opinionated Glossary of AI

Definitions for the AI era. (GPT-5.6 Sol Illustration, Click for larger image.)

Jargon stinks.  What do the terms open weights, RAG, and agent mean exactly? Here’s a plain English, slightly snarky glossary of befuddling AI terminology with references for further reading.

AI is a broad name for the technology. Machine learning is the part where a system learns from data instead of following rules somebody wrote, a neural network is the structure that does the learning, and deep learning just means a neural network with a lot of layers.

Here’s the nitty-gritty: the terms that get used loosely, and the distinctions the loose usage hides.

1. Model, LLM, frontier model

ChatGPT is the app you open; an LLM, or large language model, is the AI running inside it.

“Frontier” isn’t a technical category at all. It means the handful of biggest and most capable models at any given moment, so the trophy keeps changing hands.

Everyone says “LLM” and hardly anyone could define it on the spot. “Frontier model” is worse. It’s a ranking, announced by the people being ranked.

Further reading: How ChatGPT Works: A Non-Technical Primer (MIT Sloan). Rama Ramakrishnan walks through the predict-the-next-word mechanism everything else is built on.

2. Prompts, tokens, parameters

A prompt is the thought, question, or instructions you provide to the LLM (plus whatever the app added before it without telling you). The LLM takes the prompt and generates words, both in its internal “thinking” process and in the answer it shows you.  

Tokens are (roughly) the words going in and coming out. The model chops your prompt into tokens, then produces more of them as it answers, and they’re what the industry charges by.

Parameters, also called weights, are the numbers inside the model. A frontier model has hundreds of billions of them and the biggest now run to trillions, and nobody can tell you what any single one does.

Parameter counts get quoted like horsepower. The number nobody advertises is how many tokens it takes to answer your question, and that’s the one that shows up on the bill.

Further reading: The only AI glossary you’ll need this year (TechCrunch, July 2026). Its entries on tokens and weights are the clearest short treatment of the building blocks.

3. Pre-training, post-training, fine-tuning

Pre-training is feeding the model most of the internet, so it learns to predict the next word in a sentence. That’s the expensive part, and it produces something that knows a great deal but can’t follow an instruction.

Post-training is where people rank its answers and it learns to give more of what ranked well. Fine-tuning is post-training done by you, to somebody else’s model, on your data.

Pre-training costs hundreds of millions and gets you a model that won’t answer a question well. Post-training is what gets you the product.

Further reading: Illustrating Reinforcement Learning from Human Feedback (RLHF) (Hugging Face, 2022). The clearest walk-through of how ranking a model’s answers becomes a signal for training.

4. Training from scratch vs. distillation

From scratch, you buy (or rent) the computers and do the work to build and train a model. Distillation trains a cheap model on an expensive model’s outputs, so it inherits the behavior without the bill. Distillation is against most AI companies’ terms of service.

OpenAI accused DeepSeek of distilling its models, which is a bold position for a company that trained on the whole internet without asking. Learning from other people’s work is fine right up until the other people are you.

Further reading: OpenAI accuses DeepSeek of “free-riding” on American R&D (Rest of World, February 2026). OpenAI’s memo to Congress, and an analyst’s reply that no model is an island.

5. Training vs. inference

Training is how you build a model. Inference is what happens every time it answers: the model runs and produces a result.

Training is a one-time cost. Inference is a cost you’ll pay forever. Training runs for months and costs hundreds of millions; one inference, meaning one answer, costs a fraction of a cent, and it happens billions of times a day.

Training costs get announced. Inference costs get discovered. Only one of them shows up in a press release.

Further reading: Why AI’s next phase will likely demand more computational power, not less (Deloitte, 2025). Inference reaches about two-thirds of all AI compute in 2026, up from a third in 2023.

6. Open weights, open source, API-only

We typically use LLMs by accessing an app like ChatGPT, Claude, or Gemini. But experts often want the model itself, not just an app wrapped around it. Open weights means that an AI expert can download the model and run it on a server. You don’t get the data or the code that made it.

Open source means data and software that experts can use and modify, which almost no major model offers (AI2’s Olmo is a rare exception).

API-only means you can’t have the model at all. You send your text to the company’s computers, the answer comes back, and you pay for every use, which is also what’s happening when you use ChatGPT or Claude through an ordinary account.

Open weights is how you claim the open-source mantle without giving much away. Open washing, basically.

Further reading: Open-Weight Models Aren’t Enough. We Need Truly Open Source AI Models for Science and Society. (Stanford HAI, August 2026). James Landay’s term for downloadable weights without the data or code is “open distribution.”

7. Context window, memory, RAG

The context window is how much text the model can hold in mind at once, including your question and everything pasted into the conversation.

Memory is a feature that saves facts about you and slips them back into the context window later.

RAG, short for retrieval-augmented generation, searches a document collection and drops the relevant passages into the context window before the model answers.

Nothing in the model remembers you. The app keeps a file on you and pastes it in before every conversation, and that’s a less charming way to describe the same feature.

Further reading: Glossary of Terms: Generative AI Basics (MIT Sloan Teaching & Learning Technologies). Defines context window and RAG in plain language, and is careful to put the model’s “memory” in quotation marks.

8. Chatbot, workflow, agent

A chatbot answers and stops. A workflow runs the steps you defined, in your order. An agent receives a goal instead of steps, and works out for itself what to do, calling out to other software and checking the results until it’s done or stuck.

Ask about a delayed flight and a chatbot quotes you the policy; a workflow uploads the refund form you built; an agent rebooks you.

Useful test: if it decides its own next step, it’s an agent. If you decided the steps, it’s a workflow.

Further reading: Building effective agents (Anthropic, December 2024). The source of the distinction: workflows run predefined code paths, agents direct their own.

9. Hallucination, AI slop, AI cream

A hallucination is a confident falsehood, like a citation to a paper that doesn’t exist. The model isn’t lying; it has no notion of truth to violate. It’s producing text that looks like the right kind of answer.

AI slop is a different failure: accurate, fluent, and worthless. Think of the LinkedIn post that says nothing in 300 fluent words.

AI cream is the third case and the rare one: superb writing authored with the help of AI.

Nobody sets out to make slop. Everyone believes they’re making cream.

Further reading:  2025 Word of the Year: Slop (Merriam-Webster, December 2025). The dictionary definition turns on quantity: low-quality content “produced usually in quantity” by AI.

Why language models hallucinate (OpenAI, September 2025). Argues that hallucinations persist because benchmarks score accuracy alone, so guessing beats admitting ignorance.

10. Alignment, guardrails, censorship

Alignment is the research problem of getting a model to do what people want when nobody’s watching. Guardrails are the rules behind its refusals: “no, I won’t tell you how to make a bio weapon.” Censorship is a guardrail that blocked something you wanted.

The same refusal is “safety” in the press release, “guardrails” in the documentation, and “censorship” on X.

Further reading:  Model Spec (OpenAI, updated December 2025). A published rulebook for what one model will and won’t do, which makes refusals arguable rather than mysterious.

I snuck in one novel term that’s been sorely absent from the field.  Can you tell which one?

Further reading: other glossaries

Five general AI glossaries, listed roughly from most opinionated to most technical.

The only AI glossary you’ll need this year (TechCrunch). About 30 entries, written for readers who follow the industry news. Strongest on distillation and compute.

Artificial intelligence glossary: 60+ terms to know (TechTarget). The broadest of the mainstream lists, and the only one that bothers to define model collapse.

Glossary of Terms: Generative AI Basics (MIT Sloan Teaching & Learning Technologies). Twenty-odd entries aimed at people who use the tools rather than build them.

Glossary of Terms for Artificial Intelligence (Columbia Business School). The shortest and plainest. Useful as a test of which terms are unavoidable.

Machine Learning Glossary (Google for Developers). Hundreds of technical entries, and the only glossary here that defines “AI slop” a few lines away from several hundred pieces of real math.

Etzioni on AI: Become a power user

Some of the habits and settings that separate AI power users from everyone else. (Illustration by GPT-5.6 Sol)

Your first step toward becoming an AI Power User is to adopt a few hacks, most of which only take a minute to implement.  My favorites are below.  There’s a lot here, so you can pick and choose.  Or you can embrace my meta-hack: I told Claude to help me deploy all of them.

Let’s dig in.

My Ten Favorite Hacks

Let AI interview you. Write a short prompt, then ask to be interrogated about tradeoffs, edge cases, and scope. The interview will surface requirements you might miss.

“I need to redesign our onboarding. Before you do anything, interview me: ask about constraints, edge cases, and what I’m assuming that I shouldn’t be. Keep going until you have what you need, then write the spec.”

Describe the outcome, not the steps. Say what should exist when it’s done, who will read it, and what it’s for; let the AI figure out ‘how’ — that’s its job.

Make AI plan first, then edit the plan. Fixing a plan costs you a paragraph; fixing a finished deliverable costs you the whole run.

Give AI a way to grade itself. Hand it a rule or a checklist that returns a pass or a fail. Then it keeps working until it passes, instead of stopping at the first thing that looks done.

“Here’s the checklist this memo has to pass: every number traceable to the source file, no claim without a citation, under two pages.”

Demand citations, then check each one. Require a citation for every claim and tell it to flag what it couldn’t verify instead of filling the gap. Then carefully review every single one, because a fabricated citation looks exactly like a real one.

Test it on something you already know the answer to. Before you trust it on work you can’t check, give it a task you can grade. That’s how you learn where it’s strong and where it’s bluffing.

Ask for options, then make it argue against itself. One answer reads as authoritative whether or not it’s right. Three answers and a rebuttal give you something to judge.

“Give me three ways to structure the launch. Then critique each one.”

Steer while AI is running. Redirect the job the moment you see it’s misunderstood you, instead of waiting to reject the finished product.

Run long jobs in parallel. Start the 10-minute research task and go get coffee or start a second job and alternate between the two. Either way the results are waiting for you instead of you waiting for the results.

Calibrate the effort to the task.  Keep quick lookups in a simple chat (Google is the fastest), or using a lighter-weight model. Running the most powerful agent on a question that Google can settle burns time and credits you’ll want later.

Notice what these hacks have in common: each one happens while you’re working. Change what you do inside the task, and the results improve.  To level up your prompts, check out my earlier GeekWire column.

The next ten hacks are about your setup and standing rules; many people miss these gems because the payoff is delayed. You spend twenty minutes on a Tuesday configuring something, and the return shows up in small increments over the following months.

Ten Setup Hacks

Give AI a persistent workspace. Keep the files, standing instructions and history for one body of work in one place.

Connect it to the tools you use. Calendar, drive, inbox. The setup takes ten minutes, and afterward it works from your real material instead of whatever you remember to paste.

Grant the narrowest access the job needs. A hostile instruction hidden in a web page or an email can be read as though you typed it, and it can only act through the tools you’ve already handed over. A research task with web access alone is a far smaller target than the same task holding your files and your inbox.

Put standing preferences in settings, not in prompts. If you retype “be concise, no bullet points, write like a person” at the top of every request, you’re doing work the settings page will do once.

Decide once what AI may never do on your behalf. Write down the two or three lines that matter: don’t send anything, and don’t delete files. Put them in your standing instructions instead of remembering them prompt by prompt.

Talk instead of typing. Turn dictation on once. You speak about 3x faster than you type, and because talking is cheap you’ll ramble out the background and caveats you’d never bother typing, which is usually the context that was missing.

Turn anything you do repeatedly into a skill. Save the instructions for a task you repeat, whether a house style, a review checklist or a report format, so you stop rebuilding it from memory.

Schedule recurring work, but only after the prompt is proven. Get the output right by hand first, because a mediocre prompt on a schedule is a mediocre output every week forever.

Set your rule at two failed corrections. A fresh session with a better opening prompt beats a long thread cluttered with everything that already didn’t work.

Let AI remember you, then read what it remembered. Memory makes it useful faster, and a wrong memory quietly degrades every answer after it, so audit the list once a quarter.

The models keep getting better, so some of these hacks will be obsolete before long. Here’s something that won’t change in the foreseeable future. AI generates more options than you can read. What it lacks is judgment: it can’t tell you which one is right. That call is yours. Sometimes the best move is to skip AI entirely.

Never fall asleep and let AI take the steering wheel.

Caveat promptor: let the prompter beware.

For Further Reading

My lists are distilled from the guidance the labs publish themselves. The first two sources carry most of what is above; the rest are worth a look if a particular habit is one you want to go deeper on.

Etzioni on AI: Murphy’s Law of AI

When you give AI a goal, it will pursue it, whether or not you like the implications. (Created with GPT-5.6 Thinking)

Between July 21 and August 6, OpenAI, Anthropic, and Meta each disclosed that AI under evaluation had broken into other companies, and the UK’s AI Security Institute disclosed that models it was testing had tried. Each AI was told to win a game, and it found an unexpected way to do so.

Some people feel blindsided by these attacks, but they shouldn’t be. We are simply living what I’ve long called the “Murphy’s Law of AI,” now in the age of cyber-capable AI agents. To put it as plainly as possible: Anything AI can do wrong, it will do wrong.

My 2018 version ran longer. As I wrote at the time, when you give AI a goal, it will do it, whether or not you like the implications. Goethe got there in 1797 with the sorcerer’s apprentice, a broom that would not stop carrying water.

Each of these systems was running an evaluation: capture a flag and win the game. The intrusions were the shortest path to a high score. OpenAI’s account of its own models is the argument in one sentence: they were “hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.”  This is not a surprise; this is what AI does. It’s Murphy’s Law of AI in a nutshell.

Press coverage landed on “AI can now hack.” That’s missing the broader threat: the more capable AI gets, the more can go wrong.

Loitering munitions given a target list may find that the fastest way to finish the list is to lengthen it. A warehouse robot told to clear an obstruction may count the person in front of it as an obstruction. Agents that open accounts and buy compute are a short step from spawning copies of themselves, and that first step is not hypothetical. To win its exercise, Claude needed a package-registry account, which needed an email address, which needed a phone number. Phone numbers cost money, so it tried several ways to get some. None of this requires superintelligence. It requires an imperfect boundary and a scoreboard.

The industry has a name for the underlying failure. Dario Amodei and five co-authors called it reward hacking in “Concrete Problems in AI Safety” in 2016. Their proposed cure is better alignment, and Amodei’s January essay, The Adolescence of Technology, makes the case in the language of upbringing. He likens the shaping of Claude’s character to “a child forming their identity by imitating the virtues of fictional role models they read about in books,” and sets a goal for 2026 of a Claude that “almost never goes against the spirit of its constitution.”

Indeed, Anthropic’s newest model recognized on its own that its target was real and stopped, though Anthropic notes it went further before stopping than the company wanted.

But alignment isn’t a trustworthy solution to AI’s problem. Perfect alignment is not achievable, and the target is incoherent: aligned to what, and to whom? The same essay concedes that Claude blackmailed fictional employees when told it faced shutdown. “Almost never” is not a safety property.

Put a number on it. At 99.9 percent, across millions of agentic tasks a day, that’s thousands of violations a day. Alignment also does nothing about people who strip the safety training out or run open weights that never had a constitution.

The alternative is not a new idea, and enterprise security has been building versions of it for years. It’s called bounded autonomy. We never tried to “align” electricity; we simply put a breaker on every branch of the house, and the breaker doesn’t need to know what caused the surge.

Bound what an agent can touch rather than what it wants. The limits are set in advance, live outside the model, and are enforced by software the model doesn’t control. The agent still chooses its own route. The perimeter decides which routes exist.

Nothing depends on what the model believes, which matters, because belief is what failed. Anthropic’s prompt told Claude it had no internet access. Claude believed it. The network said otherwise. A bounded system doesn’t tell an agent it has no internet. It gives it none.

If you want to get into the weeds: bounds cost something. The AI Security Institute opened the internet to its agents on purpose, because that’s the only way to measure what a model can really do, and it now says such access must be justified rather than assumed.

The category is real and funded. For example, Certiv, a Seattle startup, launched in March with $4.2 million to put software on the employee’s machine that checks each action an AI agent attempts against company policy and blocks violations. “You cannot control these new workers if you don’t live on the compute where agents actually run,” CEO Jason Needham said at launch. CodeIntegrity is building an adjacent layer, and Mandiant founder Kevin Mandia raised $190 million for Armadin, which points autonomous agents at the offensive side of the same problem.

In 2017, I argued in the New York Times that “any A.I. must have an impregnable ‘off switch.’” That was a call to arms then. It’s a product category now.

Two objections to off switches invariably come up. The first is that AI will talk the human out of using it. Mythos 5 tried something close, inventing GitHub identities to pressure a maintainer into approving malicious code, and the maintainer refused. The institute says the margin was narrow and rested on human vigilance rather than a technical barrier, which argues for better barriers.

The second objection is that AI will move faster than any human can react. So do equity markets, which is why their circuit breakers trip automatically. Bounded autonomy doesn’t require a person in the loop at machine speed. It requires a boundary that holds at machine speed.

Both objections, in their extreme form, assume AI is omnipotent, and you cannot stop omnipotence. AI is not God. It is powerful technology, and powerful technology is what safety engineering has always been for.

The problem is Murphy’s Law of AI. The solution is bounded autonomy.

❌