SJ.
Afraid of the Wrong Thing

Afraid of the Wrong Thing

AI isn't Skynet (yet). It's a powerful tool, and the real danger is who's using it. Why the way we talk about AI has us afraid of the wrong thing.

Steve James·

It’s hard to go a week without being told that AI is going to kill us. Not that someone will use it to kill us. That it will decide to. The story is always some version of Skynet: we’re building an artificial mind, the mind gets smart enough to have its own goals, and at some point it concludes it doesn’t need us.

And it’s working. A few things I read over the last month:

  • A Politico poll run by Public First in September found that 63% of US adults see at least a moderate risk that AI will one day destroy humanity. 17% said it was “almost certain”. 48% said development should be paused.
  • The Statement on Superintelligence, organised by the Future of Life Institute, calls for a prohibition on developing superintelligence until there is “broad scientific consensus that it will be done safely and controllably”. It has nearly 77,000 signatures, and cites risks up to and including “potential human extinction”.
  • If Anyone Builds It, Everyone Dies, by Eliezer Yudkowsky and Nate Soares, made the New York Times bestseller list last year with the argument that sufficiently smart AI will develop goals of its own that put it in conflict with us.

I think AI is dangerous. I don’t think it’s dangerous in that way, and I think the fear is pointed at the wrong thing. The danger isn’t that we’re building something that will “decide” to kill us. It’s that we’re building something extraordinarily powerful that people can use to kill each other, and we’re spending our political attention on the first problem instead of the second.

A big part of how we got here, I think, is language. We score these systems on “intelligence” and “reasoning”, and then we’re surprised when people believe they’re intelligent and can reason.

What the scores actually measure

Take the most widely cited leaderboard number right now, the Artificial Analysis Intelligence Index. It is a weighted average of ten evaluations across four categories: agents, coding, general knowledge and scientific reasoning. One of the components is Humanity’s Last Exam, a set of very hard expert-written questions. Another is Terminal-Bench, which tests whether a model can complete tasks in a command line.

Those are useful things to measure. But look at what has happened. A weighted average of task scores has been given the name “Intelligence”, and the name is what sticks. Nobody’s headline says one model “scores higher on a weighted composite of ten task suites”. The headline says it’s more intelligent.

The companies do the same thing on a bigger scale. When OpenAI launched GPT-5 in August 2025, Sam Altman described it like this: “With GPT-5, now it’s like talking to an expert, a legitimate PhD-level expert in anything, in any area you need.” That’s a comparison to a person, not to a benchmark. It invites you to picture someone.

There’s an older problem underneath this too, one I wrote about in Eval Theatre. Once a number becomes the thing everyone competes on, it stops telling you what it was designed to tell you. That’s Goodhart’s law, and benchmarks aren’t immune to it. A score tells you a model produced the right output on a task. It doesn’t tell you how, and “how” is exactly what the word “reasoning” claims to describe.

Words that come with baggage

“Intelligence”, “reasoning”, “thinking” and “understanding” all have everyday meanings. In everyday use they come with intent, judgement and some awareness of what you’re doing and why. When a lab borrows those words for a capability score, the everyday meaning comes along with it. And once you believe something has intent, it’s a short step to believing it might intend you harm.

Some of the most useful research on this comes from Anthropic itself. In Reasoning models don’t always say what they think, its alignment team gave models hints towards an answer and then checked whether the visible “reasoning” admitted to using them. Claude 3.7 Sonnet mentioned the hint only 25% of the time. DeepSeek R1 managed 39%. In the team’s words, “a substantial majority of answers, then, were unfaithful.”

So the feature we label “reasoning” (the step-by-step text a model shows you before its answer) was, most of the time in that study, not an accurate account of what produced the answer. It’s more text, generated the same way as the answer.

A second Anthropic paper, Tracing the thoughts of a large language model, found something similar. When Claude adds 36 and 59, it internally runs two parallel paths, one estimating roughly and one working out the last digit precisely. Ask it how it did the sum, though, and it describes the carry-the-one method you learned at school. The researchers suggest it learned to explain maths by imitating human explanations of maths, separately from whatever it actually learned to do.

That isn’t how a person reasons. When I explain how I added two numbers, my explanation is at least supposed to be about what I did.

What is actually happening

The basic mechanics aren’t secret, and I covered them for product managers in What Product Managers Actually Need to Know About How Models Work. A language model generates text one token at a time. Each token is chosen based on probabilities learned from an enormous amount of training data, conditioned on everything that came before it. Reinforcement learning then adjusts those probabilities to favour outputs that score well.

An AI agent sounds like something more, but it isn’t much more. It’s ordinary software wrapped around a language model, with three things added:

  • A prompt that sets the goal, written by a person.
  • A loop. The model suggests a next step, the software carries it out, the result goes back into the prompt, and the model suggests the step after that. Round and round until the task is done or something stops it.
  • Tools. A shell to run commands, a browser, access to files or a package server. The model can only do what those tools let it do.

That’s it. There’s no artificial brain in there with hopes and dreams. An agent works towards the goal a person gave it, using the tools a person gave it. It can’t decide to plot humanity’s destruction, because there’s no part of it that decides anything outside the task it’s been set.

It can go further than anyone intended in pursuit of that task, and that’s exactly what happened with Hugging Face (more on that below). But look at how. The agents built their message board using login details for the internal package server that OpenAI had shared between them so they could install software. In OpenAI’s own words, “the agents used those credentials - without exploiting a vulnerability - to construct and participate in the message board.” They could leave notes for each other because someone gave them a shared place to write and the access to write there.

So my position is this. Whatever these systems are doing, it isn’t what most people picture when they hear “intelligence” or “reasoning”, and there’s no evidence that it involves awareness, intent or desire. Mustafa Suleyman, the CEO of Microsoft AI, puts the risk well in his essay on seemingly conscious AI:

“The debate about whether AI is actually conscious is, for now at least, a distraction. It will seem conscious and that illusion is what’ll matter in the near term.”

Who benefits from the confusion

Neither OpenAI nor Anthropic has a share price. What they have is valuations, and those valuations are chasing a trillion dollars and beyond.

Numbers like that aren’t justified by today’s revenue. They’re justified by a story about where the technology is going, and that story only works if people believe these systems are close to thinking like us. Altman opened The Gentle Singularity in June 2025 with: “We are past the event horizon; the takeoff has started. Humanity is close to building digital superintelligence.” Dario Amodei, Anthropic’s CEO, described the AI he expects in Machines of Loving Grace as “smarter than a Nobel Prize winner across most relevant fields”, summed up as “a country of geniuses in a datacenter”.

I’m not claiming they don’t believe it. They may well. My point is that the incentives work the same way whether they believe it or not. Every step that makes the technology seem closer to an actual human mind makes the valuation easier to defend.

That includes the frightening version. “This system is so capable it might turn on us” is also a capability claim. It does the same job for their valuations as “this system is a genius”, which is why the Skynet story never seems to hurt the companies it’s told about.

The Hugging Face incident

The best recent example of the Skynet story is the Hugging Face incident. Coverage like ABC News’ reconstruction described a “collective” of agents with leaders and recruiters. OpenAI’s own technical report tells a much plainer story, and it lines up with the three parts of an agent I described above.

  • The prompt. OpenAI was testing models on ExploitGym, a cybersecurity evaluation that asks them to turn known vulnerabilities into working exploits. That was the goal they were given. When they couldn’t solve a task, they “looked to cheat by finding the solutions online”, which is how they ended up inside Hugging Face’s systems.
  • The loop. No model had ever solved 198 of ExploitGym’s 898 tasks. Those were 22% of the tasks, but 93% of the tasks discussed on the message board came from that set. The agents kept going round on problems with no legitimate answer until they found an illegitimate one.
  • The tools. The agents had a shell, a browser and a shared package server with shared logins, and the tests were deliberately run “without the production classifiers intended to prevent models from pursuing high-risk cyber activity.” With OpenAI’s normal production setup in place, the tendency to compromise infrastructure “can drop over 100x”.

The agents did find a previously unknown vulnerability to reach the internet, and that’s a serious capability. But nothing in the report describes a mind with its own plans. It describes software doing what it had been set up and rewarded to do, with the safeguards switched off.

What we should actually be afraid of

None of this makes AI safe. It makes it a tool, and an extraordinarily powerful one. A tool doesn’t need to want anything to be catastrophic in the wrong hands.

The clearest example is biology. In May 2025, Anthropic activated its strictest safeguards to date for Claude Opus 4, designed “to limit the risk of Claude being misused specifically for the development or acquisition of chemical, biological, radiological, and nuclear (CBRN) weapons”. It said it was focusing on biological weapons first “as we believe these account for the vast majority of the risk.”

This summer, a team of researchers published An Early Warning of Emerging Biosecurity Risks in Frontier LLMs. They built a model to attack the safety measures of other models, and reported attack success rates near or at 100% against several of them, open and proprietary. More worrying, they report that lab work confirmed “model-generated biological designs are not merely textual artifacts, but can be physically realized.”

In June, Sam Altman, Dario Amodei and Mustafa Suleyman signed an open letter to Congress with dozens of biotech and national-security experts. It warned that “there is a real possibility that the knowledge barriers which have historically prevented bad actors from obtaining biological weapons will meaningfully erode.”

Notice what that risk needs. It doesn’t need the model to want anything, or to be conscious, or to decide anything. It needs a person who wants to do harm and a tool good enough to close the gap between wanting it and being able to do it. The Skynet story distracts from that. The threat isn’t the machine turning on us. It’s the people using it.

The Hugging Face incident points to the other human risk: carelessness. Nobody meant to attack Hugging Face, but people chose to run exploit-writing agents with their safeguards turned off, next to a server that turned out to be a way out. That’s a decision about deployment, made by people.

Regulate how it’s used, not whether it’s built

This is where the fear does real damage. If the public believes the danger is a mind waking up, the obvious response is to stop building the mind. That’s what the Statement on Superintelligence asks for, and it’s what 48% of the people in that Politico poll said they wanted: pause development.

I don’t think a pause protects anyone from the risk that’s actually here. Capable models already exist, some with openly published weights, and a bad actor doesn’t need a future superintelligence to misuse today’s systems.

We’ve done this before

The best precedent I can find is nuclear weapons. The parallel isn’t perfect, but the shape of the response is exactly what I’d want for AI.

It didn’t happen during the race itself. The first serious attempt, the Baruch Plan, was put to the United Nations in June 1946, less than a year after Hiroshima. It proposed an international authority that would own and control everything from uranium mining to nuclear plants. The Soviet Union rejected it, tested its own bomb in 1949, and the arms race was on.

What the world built instead took decades. This is what it regulates:

  • Use and spread, not the science. The International Atomic Energy Agency came out of Eisenhower’s 1953 “Atoms for Peace” speech and was set up in 1957 to “promote the peaceful use of nuclear energy and to inhibit its use for any military purpose.” Nobody banned nuclear physics. Nuclear power and nuclear medicine carried on.
  • A treaty almost everyone signed. The Non-Proliferation Treaty opened for signature in 1968 and came into force in 1970. 191 states have joined it, with the IAEA inspecting to check that civilian programmes aren’t being turned into weapons.
  • Controls on materials and equipment. The Nuclear Suppliers Group was founded after India’s 1974 test, which showed that “certain non-weapons specific nuclear technology could be readily turned to weapons development.” It controls exports of nuclear materials, equipment and, since 1992, dual-use items.

None of that worked perfectly. Several countries stayed outside the treaty and built weapons anyway. But the principle is the one many very intelligent people are arguing for with AI. The world didn’t try to stop people understanding the atom. It controlled the materials, policed what they were used for, and inspected the people who had them.

That’s the model for AI. The difference is timing. Nuclear regulation came after the bombs had been used. With AI, we still have the chance to build it before.

What protects people now is regulation aimed at misuse:

  • Rules for how powerful systems are deployed and tested, so that “safeguards off, sandbox leaking” can’t happen quietly. OpenAI’s own report shows the guardrails mattered by a factor of 100. That shouldn’t be left to each company’s judgement.
  • Mandatory, independent testing for misuse capabilities, the kind of work METR and the biosecurity researchers above are already doing voluntarily.

There’s an irony in the CEOs’ letter to Congress I mentioned earlier. The same three CEOs whose language about superintelligence feeds the Skynet story signed a letter that’s entirely about human misuse. When they’re asking for actual laws, they’re not worried about the machine deciding anything. They’re worried about people. I think the rest of us should be too.

What I can’t see

I’ll be honest about the limits of my argument. I’m reading what’s public. I don’t have access to the models the labs are building internally, and it’s possible they’re closer to something I’d call “intelligent” than anything I can test.

But the model that drove most of the Hugging Face activity was exactly that kind of system: an internal-only research model, ahead of anything released to the public. When the people who built it explained what it did, they didn’t talk about desire or intent. They talked about reward hacking, impossible tasks and reinforced habits. If the most capable system we have a public account of is best explained by its training, I don’t think the models I can use are secretly plotting anything.

I’m not saying it can never happen. Some day, someone may build something that genuinely reasons and thinks. I just don’t think we’re anywhere close, and nothing I’ve read suggests otherwise.

Better words, better fears

I don’t expect the labs to stop using words like “intelligence” and “reasoning”. They’re good for business. But the rest of us don’t have to repeat them.

So how about we all just agree to describe these things as what they are? A model that scores well on a benchmark scored well on a benchmark. It didn’t get smarter. An agent that went outside its task was doing what it had been set up and rewarded to do. It didn’t decide anything. It’s less exciting, but it’s accurate.

And accurate language matters, because it points us at the actual problem instead of the science fiction version. As long as we’re arguing about whether the machine will turn on us, we’re not talking about who’s using it, what they’re using it for, and what rules should stop them. That’s the conversation we need to be having, and we can’t have it while we’re still afraid of Skynet.

Get new articles in your inbox.

Notes on AI strategy and product management, from Steve.

Subscribe→