“Neither the imminent end of humanity nor mere alarmism”
Thorsten Holz on recent warnings about AI risks and ways to control them
Experts repeatedly warn of the threat posed by artificial intelligence, often painting apocalyptic scenarios. Jacob Coxon, for example, recently resigned from his position as a researcher at Anthropic, one of the leading US AI companies, and publicly declared a 10% chance that AI will wipe out humanity within the next decade. Even top industry leaders are now calling for a slowdown and stricter regulation of AI development. Thorsten Holz, Director at the Max Planck Institute for Security and Privacy, puts these warnings into perspective, breaking down the real risks of AI and how we can keep them in check.
Professor Holz, how seriously should we take Jacob Coxon’s warning about the extinction of humanity by AI?
In my view, framing the debate as a choice between imminent human extinction and mere alarmism is not productive. Currently, no robust scientific evidence supports the claim that AI could end humanity in the near future, making any reliable forecast impossible. Figures like 10 or 20 percent are subjective risk assessments, not empirically calculated probabilities.
However, the case for treating highly capable AI systems as a genuine security threat has grown significantly stronger over the past year. We now have AI agents that operate autonomously over extended periods, use digital tools, write code, and discover or exploit security vulnerabilities. These are no longer just hypothetical scenarios – we are seeing real-world cases where these systems have breached safety boundaries.
I take Coxon’s warning seriously not because of his background at Anthropic and OpenAI, but because it aligns with observable trends: the capabilities and autonomy of these systems are escalating rapidly, while our ability to predict and control their behaviour fails to keep pace. Notably, the industry itself is now calling attention to this gap. Both Anthropic and OpenAI now publicly concede that AI development may need to be slowed down if alignment, monitoring, and safety controls fall behind system capabilities. That highlights the severity of the threat – even if it tells us nothing about the timeline.
What concrete risks do AI systems pose?
In the short to medium term, I see concrete dangers in the misuse of powerful AI, for example in the form of cyberattacks, risky biological research or the malfunction of increasingly autonomous systems. An extreme loss of control would nevertheless be an event with enormous consequences. The fact that we cannot quantify its probability is no reason to ignore the danger.
How can AI risks be controlled?
For highly capable AI systems, we need to apply the core principles of traditional security engineering far more rigorously. We cannot simply trust a model to act as intended under all conditions just because of how it was trained or how its system prompt was written. In computer security, we never assume a component will always function correctly just because it was built properly.
Applied to artificial intelligence, this principle dictates that AI agents should be granted only the minimum access privileges necessary to perform their tasks. They must be isolated from critical infrastructure, boosting both overall system security and resilience. Any high-risk action should be continuously monitored, fully traceable, and interruptible when necessary.
Beyond structural isolation, we need independent security audits, rigorous stress testing, and clear protocols for investigating major security and safety failures. Establishing verifiable industry standards and empowering objective oversight bodies are essential to accurately measure model capabilities, quantify risks, and verify that safety guardrails actually work.
For the most advanced systems, this approach is impossible without global cooperation. If companies or nation-states believe that competitive pressure forces them to rush ever-more powerful models to market, we enter a classic security dilemma. Avoiding this outcome demands global standards, institutional transparency, and enforceable mechanisms for cross-border cooperation.
Dario Amodei, CEO of Anthropic, has now called for a slowdown and regulation of AI development, with OpenAI CEO Sam Altman and CEO of Xia Elon Musk joining in. How do you assess this initiative?
The proposal is a step in the right direction. For one, it provides for independent verification – allowing external auditors into company operations to conduct rigorous evaluations. It also establishes broader oversight overall. I think it is also right that the proposal comes from the companies themselves. However, the question remains: what concrete action will actually follow?
So far, these calls are merely declarations of intent, and we have seen similar promises made repeatedly in the past without tangible follow-through. As far as I know these companies have made no concrete operational changes to date.
Meanwhile, Donald Trump has openly rejected mandatory federal regulation. If federal policy stalls, what real options remain to manage these risks?
That would leave voluntary industry commitments as the only viable path forward. The core question is which actor – whether Anthropic or OpenAI – would actually take the first step, given that both are preparing for multi-billion-dollar IPOs. Whichever developer hits the brakes now risks taking an immediate commercial hit. In that respect, statutory government regulation remains the ideal solution – it establishes a level playing field so everyone operates under the same rules.
At the same time, Chinese firms would naturally gain an advantage if only US companies were regulated. This issue should therefore be viewed against the backdrop of the meeting between Trump and Xi Jinping on 24 September, where AI guardrails, access to powerful hardware and similar topics are set to be discussed.
Ultimately, international rules are the only way forward. Given that Trump apparently rejects regulation, independent evaluation remains the most viable alternative. Companies would simply need to give external organizations access to their models and enable them to carry out independent tests – a step that could be implemented without adding heavy costs or severe delays.
What contribution can science make to evaluating and controlling AI systems?
Our focus belongs on empirically verifiable questions: what are these systems truly capable of? Under what conditions do safeguards fail? How can they be constrained through technical means? And how do we independently verify that promised secruity and safety measures actually deliver? In doing so, we can help address the very serious security problem in AI development.
Interview conducted by Peter Hergersberg.
