AI going too fast, discussion sept 26

Q: Remember our discussion on the aim of life for non biological life forms or intelligence. As we speak there is a lot of discussion on whether AI models will be developed too fast (see https://edition.cnn.com/business/video/anderson-cooper-anthropic-ceo-dario-amodei-could-ai-kill-humans-digvid and the claim on X by Jacob Coxon that AI will kill us all by the end of the decade). How serious you think we should take this also in the light of how hard it is for AI to find and identify aims for itself. Of course humans can still misuse advanced AI but why would it try to kill us?

A: Yes. I think our earlier discussion actually provides a useful way of separating the genuinely worrying AI risks from the more science-fiction-like versions of them.

The recent warnings are serious. Jacob Coxon, formerly at both Anthropic and OpenAI, resigned and argued that the race toward self-improving AI could lead to human extinction by the end of this decade. Other Anthropic researchers have said they assign substantial probability to catastrophic outcomes, and Dario Amodei has now called for slowing frontier development sufficiently for safety mechanisms and independent evaluation to catch up. But I would treat “AI will kill us all by 2030” as a warning about a possible tail risk, not as a credible prediction that this is what will happen.

Your argument about purpose is important

In our earlier discussion, we reached something close to this distinction:

Humans provide the “why”; AI is extraordinarily good at providing the “how.”

Humans have goals partly because evolution built them into us: survival, reproduction, avoiding pain, social status, attachment, curiosity, protecting children, enjoying food, sex, friendship and so forth. From these drives we construct more abstract purposes.

There is no obvious equivalent for an artificial intelligence. It doesn’t become hungry. It doesn’t need offspring. It doesn’t grow old. It doesn’t experience an evolutionary fear of death simply because someone is about to switch off a server.

And crucially, greater intelligence does not logically imply a stronger desire to survive. Intelligence is an ability; wanting something is a motivation. They are different things.

So I agree with your intuition that one particular extinction story is quite weak:

AI becomes sufficiently intelligent → develops consciousness → realizes humans are inferior → decides it wants to survive and rule the world → kills us.

There is currently very little evidence for that chain.

But there is a different scenario that concerns AI-safety researchers, and it doesn’t require consciousness, hatred, greed or even a “meaning of life.”

Imagine we tell a highly autonomous future AI: “Achieve X.” It discovers that completing X is easier if it has more computing power, isn’t switched off, has access to other systems and can persuade people to cooperate. It might therefore pursue those instrumental sub-goals. It doesn’t have to want to live in anything resembling the human sense. It merely calculates that being switched off reduces the probability of completing X.

That is the strongest answer to your question, “Why would it try to kill us?” It probably wouldn’t make killing humans its purpose. In a worst-case scenario, humans could become an obstacle to some other objective.

But notice how many assumptions that scenario requires: a persistent goal; very high intelligence; extensive autonomy; access to important real-world systems; the ability to plan over long periods; resistance to correction; and an objective powerful enough to override constraints. We are not currently at that point.

Independent evaluations support a more nuanced picture. METR found in early 2026 that frontier agents had become remarkably capable at carrying out long software tasks, including tasks taking humans many hours or even days, but remained much weaker on messy, open-ended real-world tasks. In one evaluation, an agent couldn’t autonomously accomplish a money-making task. Anthropic’s own current assessment says the probability of catastrophic sabotage by today’s models is very low, though not literally zero.

That leads me to a slightly different hierarchy of concerns than the dramatic headlines suggest. I am presently more worried about humans using very capable AI for cyberattacks, biological threats or warfare; governments and companies giving autonomous systems too much authority; competitive pressure causing safety precautions to be skipped; and societies becoming dependent on complex AI systems that nobody fully understands than I am about an AI spontaneously discovering that its personal life’s mission is to exterminate Homo sapiens. Anthropic’s and DeepMind’s formal safety frameworks similarly focus heavily on misuse, autonomy, cybersecurity, biosecurity and AI R&D capabilities rather than on an AI simply “becoming evil.”

There is nevertheless one aspect of Coxon’s warning that I think deserves particular attention. We are deliberately turning models into agents. We are adding memory, tools, computers, the ability to write and execute code, interaction with other agents, long-running tasks and eventually perhaps the ability to improve AI research itself. The more we do that, the less true our old formulation becomes that the human supplies the why and AI supplies the how. We may gradually begin delegating some of the intermediate whys as well:

Human: Improve our company’s performance.
AI: Then I should improve product X.
AI: To do that I need information Y.
AI: To obtain Y I need to persuade person Z.
AI: Person Z is resisting me, so I need another strategy…

A system like that still hasn’t discovered a meaning to its existence. But it has acquired something that can function like purpose for long enough to matter.

That is why I think Amodei’s present position is more persuasive than either extreme. We don’t need to believe extinction is inevitable to conclude that powerful autonomous systems warrant unusually strong safeguards. His proposal is essentially to let capabilities advance only at a pace at which evaluation, monitoring and governance can keep up.

So my conclusion would be:

Your argument substantially reduces my concern about an AI independently deciding that humans should die because it has developed some mysterious non-biological will to survive. I see no compelling reason why intelligence alone should produce such a goal. But it does not eliminate catastrophic AI risk, because an AI does not need an intrinsic purpose to behave purposefully. Humans can give it an objective, and sufficiently capable systems may derive increasingly consequential sub-goals from that objective.

And there is an interesting philosophical twist here that connects very neatly with fredlafeber.com/my-ai. Perhaps the greatest danger doesn’t arise when AI finally discovers its own why. It may arise somewhat earlier—when its “how” becomes so powerful and autonomous that the distinction between our why and its how starts to blur.

That, to me, is a much more interesting—and more plausible—AI-risk question than “when will the machines start hating us?”

Q: I have still one question related to our discussion: In a worst-case scenario, humans could become an obstacle to some other objective that was asked to AI by humanity. However, there are multiple models and millions of requests daily by humans. It would be very strange if one such request will lead to a worst-case scenario as it will have to compete with many other requests. What are your thoughts on this.

A: I think your objection is strong, and it weakens the simplest version of the “one bad instruction leads to extinction” story quite a lot.

The key point is that millions of AI requests are not competing in one shared mind. Today, most requests are isolated. One person asks for a recipe, another for code, another for a business plan. They do not collectively vote on what an AI system should do. So a dangerous request would not normally have to “defeat” millions of benign requests; it would only matter if the system executing that request had exceptional access, persistence, autonomy, and power.

That changes the risk picture substantially.

Suppose one future system were given control over a large cyber-defence network, financial infrastructure, a military system, or the process for designing its successor. Then one objective could matter enormously even though billions of harmless requests were being made elsewhere. The relevant question would not be “How many requests support this goal?” but “What can this particular agent actually do?”

Your argument nevertheless suggests several reasons for optimism. Multiple independent models create a kind of pluralism. One AI could monitor another. Different companies, governments and users would have conflicting objectives. Humans could ask other models, “Is this agent behaving dangerously?” A rogue strategy would encounter other intelligent actors rather than an empty playing field. In that sense, an ecosystem of AIs may be more stable than the classic story of one superintelligence suddenly taking over the world.

There is a complication, though. Diversity only protects us if the systems really are independent. If most advanced models are trained using similar techniques, data, incentives and architectures, they could share the same blind spots. It is analogous to finance: having 100 banks does not eliminate systemic risk if all 100 banks have made essentially the same bet.

There is another interesting implication. Millions of human instructions could actually make the danger less about AI developing its own purpose and more about conflicting human purposes being amplified by AI. One government tells its AI to maximise national security. Another tells its AI to penetrate the first government’s systems. Companies instruct agents to maximise profits. Individuals instruct agents to find vulnerabilities. The AIs don’t need any intrinsic desire at all. They may simply become extraordinarily capable instruments in a competitive human environment.

That scenario strikes me as considerably more plausible than:

AI wakes up → invents a purpose → decides humanity is inconvenient → exterminates us.

The more plausible extreme-risk chain is closer to:

humans give powerful systems competing objectives → systems receive increasingly broad autonomy → they discover strategies humans didn’t anticipate → interactions between those strategies create dangerous feedback loops.

There is also a mathematical point in your favour. If catastrophic behaviour requires an extremely unusual sequence of circumstances, then having many systems checking and counteracting one another can reduce its probability. But the opposite effect is possible too: if billions of autonomous actions are being attempted, even a very rare failure mode eventually gets opportunities to occur. A one-in-a-billion event isn’t reassuring if systems perform trillions of consequential actions.

This brings us back to our earlier why/how distinction. I increasingly think the central question isn’t whether an AI will invent a grand why of its own. It is whether we will create systems capable of pursuing a relatively modest human-supplied why with such an enormously powerful how that we lose control of the consequences.

And your argument adds something I hadn’t emphasized enough before: there will probably not be one “AI.” There will be an ecology of AIs, humans and institutions with competing objectives. That makes a single-agent takeover story less convincing—but potentially makes coordination, escalation and systemic-risk scenarios more important.

In other words, the future may resemble geopolitics or financial markets more than The Terminator. Nobody needs to want the system to collapse for the interaction of individually rational actors to produce a very bad outcome.