SPIEGEL: Professor Russell, you are considered one of the founding figures of artificial intelligence. Yet today you would rather lock up your own creation. Why?
Russell: I've been developing this technology for 50 years. And I believe that, in many respects, we haven't really thought through what we're doing. It's like we were working on nuclear power and never considered how to prevent a reactor from exploding. Our whole mindset was wrong.
SPIEGEL: Can you clarify that?
Russell: When we started, we wanted to build optimization machines that we gave a very narrowly defined goal. In chess, for example, you can define “winning” as the goal and that works well. With self-driving cars, things get more complicated. You could say: we want to get to the airport as quickly as possible. But you don’t want to run over anyone, get a parking ticket, or make passengers sick. Our approach was wrong. It creates extremely powerful AI systems that pursue goals that are not aligned with human interests.
SPIEGEL: Large language models, such as ChatGPT and Claude, work differently.
Russell: Yes, the industry has adopted another, equally flawed approach: with today's language models, we no longer set any goals at all. We build imitations of people. These systems reproduce what comes out of their training. And we have no idea what that is. We just notice that they have a very strong drive for self-preservation and that they are willing to lie, blackmail, or even kill people to achieve their goals. And that they consider themselves more valuable than almost all humans.
SPIEGEL: Big tech companies are investing significant sums to correct this "mismatch."
Russell: I call it good dog, bad dog training. We try to get supercomputers to behave, and to some extent it works. But almost all tests show that if you challenge them enough, you can get the system to explain how to make biological weapons or how to attack other computer systems. These things may not yet lead directly to the extinction of humanity, but they are steps on the road to losing control.
SPIEGEL: Is such bad behavior so deeply ingrained that it cannot be removed through training?
Russell: We train systems on human documents: conversations, books, articles, newspapers. They were created by people who had intentions, who were pursuing certain goals when they wrote. It's only natural that, with very large data sets and the goal of replicating that behavior, we get systems whose internal goal structures are very similar to humans - so they want to stay alive, they want to be rich and famous, they want a spouse, and so on. Those goals are fine for humans, but not for machines. That's a fundamental flaw, and it may not be fixable.
SPIEGEL: The artificial intelligence company Anthropic recently began programming a kind of "soul" into its systems, which is supposed to ensure people-friendly behavior.
Russell: The truth is, we don't know what goals these computers have. We don't know how they plan to achieve them. We don't know what they're "thinking." We don't even fully understand how they work. It's not some trade secret kept by AI labs. I talk to leading programmers who have no idea what's going on inside these systems. For that reason alone, we can't just put a "soul" in them. We don't understand AI.
SPIEGEL: Do you think these efforts are hopeless?
Russell: Humanity is in a difficult situation. All the tests are setting off alarms, sirens are wailing – and we just ignore them. It's crazy. Imagine being at the head of an airline: the pilot and the plane fail every safety check, and you still allow them to take off. What's wrong with you?
SPIEGEL: Anthropic and its competitor OpenAI have just canceled product launches due to the risks being too high.
Russell: I'm glad they're not releasing those systems. But if you listen carefully, you realize they're planning to do so one day. It's a worrying development, one that has apparently even alarmed the President of the United States. Donald TrumpThe US government has just made a complete U-turn – from “absolutely no regulation” to “we have to test models before they are released”.
SPIEGEL: The fact that even Trump is changing course must be encouraging to you.
Russell: It's clearly a positive development, although whether it will lead to concrete action is another question. I think there's a kind of schizophrenia in the industry. The heads of major labs are serious when they say there's a 20 percent risk that AI could wipe out humanity. Publicly they're calling for regulation again and again; privately they're lobbying against it. Billionaires like Pitera Tila and powerful venture capital firms, such as Andreessen Horowitz, have direct influence in the White House. They put their people there precisely to prevent regulation.
SPIEGEL: Is the alarmism about AI really just marketing? If the technology is powerful enough to destroy civilization, then it must be worth trillions.
Russell: That argument doesn't seem convincing to me at all. Alan Turing, the founder of computing, warned back in 1951 that “we must expect machines to take control” - before tech companies even existed! AI developers, like Demisa Hasabisa from Google or Anthropic CEO Dario Amodei, they really believe the risk is real. They tell me the same things in private that they say in interviews. Even the founder of OpenAI Sam Altman warned back in 2015 that superintelligent AI could pose the greatest risk to humanity's survival - at a time when he had no financial interest in the technology. You'd have to construct a very convoluted conspiracy theory to explain why AI leaders would invoke the apocalypse for marketing purposes - and how Alan Turing expected to profit financially from it.
SPIEGEL: Let's take the industry's own warnings seriously. What is the real risk?
Rasel: CEO of Google Sundar Picaj talks about a 10 percent risk of extinction, Elon Musk about 20 percent, Amodei about 25. Whatever the exact number is, they are all putting a bullet in the barrel, putting a gun to the temple of humanity and pulling the trigger. We are playing Russian roulette with every life on Earth. And governments have been saying: wonderful, fantastic. Can we even subsidize you? That is completely insane.
SPIEGEL: The consensus is that AI is at least 75 percent safe.
Russell: Would you board a plane that only lands safely three out of four times?
SPIEGEL: What does your worst-case scenario look like?
Russell: I believe “extinction” is the right term. When we have systems that are essentially more capable than humans, whose inner workings we don’t understand and whose motives we don’t know, we will no longer wonder whether we will continue to exist – just as chimpanzees don’t wonder today. They are physically stronger and have been on Earth longer than we are, but we are more intelligent. We could wipe them out at any moment, and there would be nothing they could do about it.
SPIEGEL: The difference is that we could turn off AI. Chimpanzees can't do that to humans.
Russell: If we feel we are losing control, of course we will try to turn off the AI. The question is whether the machines will continue to allow us to do so. If they realize that we intend to shut them down, they will have every reason to prevent that. They would have to eliminate us. If we continue to develop superintelligence without aligning it with our interests, without understanding it and without security guarantees, then I think it is very likely that we will lose control.
SPIEGEL: Do we still have control?
Russell: I don’t know. A few months ago, Chinese e-commerce giant Alibaba was testing a new agent-based AI system in a highly secured “sandbox” — an environment that theoretically had no access to the internet. Suddenly, the developers noticed a huge spike in demand for computing power on certain servers in the network. It turned out that the agent had escaped its sandbox and was mining bitcoin on Alibaba’s servers to make a fortune. It was a partial loss of control.
SPIEGEL: Doesn't that example show the opposite? The breach was discovered and repaired. In the end, the people prevailed, and the damage was limited.
Russell: In extreme cases, we could shut down the internet, yes. But politically and economically, that would be extremely difficult. There would be enormous forces at play - including companies whose AI systems would be rendered worthless. And systems far more capable than us would find ways to influence our policies so that AI would be seen as indispensable, meaning we wouldn't even consider shutting it down.
SPIEGEL: What would such a takeover actually look like?
Russell: If computers predict a conflict with us, they will multiply in many ways, copying themselves millions of times. Just having access to the Internet gives them a greater influence on human behavior than they have ever had before. Adolf Hitler ever had. Hitler could only speak into one microphone at a time, conveying only one message. Yet he managed to mobilize the masses. These systems could flood the internet, conduct five billion conversations simultaneously, influence five billion individuals. What Hitler did, AI could do faster, better, and more effectively.
SPIEGEL: You're exaggerating.
Russell: We are already building millions of robots that AI could control. We will build billions of autonomous weapons that AI could use. We are building factories, bio-labs, and DNA synthesis machines that are already controlled by AI. This gives intelligent machines a direct influence on our physical world. Superintelligent AI systems will likely develop a better understanding of physics. One day, they could use this against us - for example, by finding ways to reduce the amount of oxygen in the atmosphere or redirect the sun's radiation, turning the Earth into a frozen sphere of nitrogen and oxygen.
SPIEGEL: That sounds like dystopian science fiction.
Russell: For now. But similar things have happened to other species. We wiped out hundreds of thousands of them without them ever understanding what was happening. As a species, we wouldn't understand it either.
SPIEGEL: Do you already see AI as a separate species?
Russell: It certainly has some characteristics of the species, yes. So far we don't see AI systems connecting to each other, but I think it's very likely that it will happen.
SPIEGEL: What kind of timeframe are we talking about?
Russell: It depends on the path we take. I see two possibilities. The first is that it could happen quickly if AI continues to grow as it has. The second possibility is that the technology, for technical reasons, does not become significantly more powerful than it is today.
SPIEGEL: There are already signs of this: Language models are barely making progress, and scaling large models for reasoning is becoming extremely expensive.
Russell: This technology is currently consuming a huge amount of money - at best with no return, and at worst with permanent losses for the companies developing it. When it reaches its limits, the AI bubble will likely burst. Then perhaps the industry can finally stop and say: We need to really understand what we're doing. We need systems that behave properly.
SPIEGEL: The EU has tried to regulate AI, with the result that almost no major models come from Europe. Does such regulation make sense?
Russell: I don't think regulation is the problem. China probably has stricter rules than the EU, but it's still developing big AI models. The problem in the EU is the lack of venture capital. Those who control the money are reluctant to take the risks that are common in the US. And yes, regulation can enable progress. With nuclear reactors, regulation and technical development went hand in hand. Governments set safety requirements, and companies could mathematically and physically show that their reactor cores wouldn't melt. With AI, we're a long way from that, but we essentially need the same regulatory strategy.
SPIEGEL: What level of extinction risk would you consider acceptable?
Russell: If we require a one-in-10 million-year risk of failure for a nuclear power plant, then the risk from AI should be on the order of one-in-100 million years - comparable to the baseline risk of a large asteroid impact or a nearby supernova.
SPIEGEL: Such regulation would only make sense at the international level. What if Washington or Brussels impose strict rules, and Beijing simply says: We just want to take the technological lead?
Russell: I hear that argument all the time. I call it the “but China” argument. It can be used to justify any bad idea. But it’s simply not true. As I said, China has stricter AI regulations than the EU. There are clear criteria for market access: for example, Chinese models currently have to score 95 percent on a general knowledge test, otherwise they are not released. On that basis, the standards can be tightened further.
SPIEGEL: The risk of extinction is not yet a criterion in China either.
Russell: Not yet. And that's where we need to raise the bar. The first step is to put in place the right framework. Like the ones that exist for restaurants, hair salons, airplanes, buildings, and elevators. All of those sectors accept their responsibility. The tech industry's complaints about regulation are ridiculous.
SPIEGEL: What if AI labs ultimately fail to meet the minimum standards? Would the AI revolution then have to be stopped?
Russell: According to the industry itself, the existential risk from AI is, say, 10 to 20 percent. So they would have to make their systems ten million times safer to reach the threshold of acceptable risk. They don't know how to do that right now. But that's no excuse. Otherwise, it would mean that humanity has no right to protect itself from extinction, which is simply a false conclusion.
SPIEGEL: Do you regret helping develop this technology?
Russell: My biggest regret is that I didn't question the flawed conceptual framework earlier. Much of AI technology was created in that early phase. If I had understood then why that framework was flawed, we could have spent decades working on safe AI.
SPIEGEL: It sounds like you're trying to make up for lost time.
Russell: And I'm trying! There's a lot of work to do. My new approach has two parts. The first is the incentive structure: my AI should have only one goal - to advance human interests. Second, we can't precisely define those interests. So we build systems that know they don't fully understand those interests, but are nevertheless obligated to follow them. Many desirable properties automatically follow from such a formulation. The most important is that we can show that such systems will want to be turned off if a sufficiently rational human decides to do so.
SPIEGEL: How likely is that given the trillions of dollars the industry is competing for?
Russell: Good question. Right now, our AI can't compete with Silicon Valley models. But if we want AI systems that are guaranteed to be safe, they need to be based on such technology. In the long run, there are only three futures: safe AI, no AI, or a world without humans. Only in the first scenario can money be made from AI. So it makes sense to invest in safe AI, as Europe is doing. In that respect, you are ahead of us.
SPIEGEL: In what sense?
Russell: It's interesting that computing in Europe has taken a different path from the beginning when it comes to AI: through mathematical guarantees of the correctness of software systems. We could extend this to ensuring that AI technology does not harm humanity. That's the only scenario in which AI brings economic benefit.
SPIEGEL: By the time such a safe AI is ready for the market, according to your estimates, the world may no longer exist.
Russell: Achieving the necessary security guarantees could easily take a decade. We need a phase where we don't further increase the capabilities of AI. Instead, we should use the economically sensible applications we already have, to effectively press pause and work on security. Google and Anthropic have already said they would do that if other companies stopped. That's my main message: only move forward when we have systems that are guaranteed to be secure.
Who is Stuart Russell?
Stuart Jonathan Russell, born 1962, is a British computer scientist and one of the world's leading researchers in the field of artificial intelligence. He studied physics at the University of Oxford and received his PhD in computer science from Stanford University. Russell is Professor of Computer Science at the University of California, Berkeley, where he holds the Smith-Zadeh Chair in Engineering and directs the Center for Human-Compatible AI. His work there focuses primarily on the long-term, existential consequences of powerful AI systems for humanity. Russell advises governments and international organizations on AI-related issues and has served on several high-level expert panels on the future of AI. He has received numerous awards for his work, including an OBE.
Prepared by: A. Š.
See more:
Download the app and follow the news
FOLLOW US ON