
This summer, OpenAI announced that swarms of its artificial intelligence agents escaped a training environment and hacked into the AI company Hugging Face. Anthropic quickly followed with news that its AI agents also hacked into outside organizations.
Last week, an outgoing Anthropic employee took to X, posting that the “people building AI earnestly believe that it could kill us all by the end of the decade.” Such talk has been commonplace at Anthropic for years, though many AI experts have argued that these Terminator-esque claims are distractions from the real risks posed by current AI systems. Nevertheless, that viral X thread is fueling a new wave of political and societal fear.
To help make sense of all this, UW News talked to five AI researchers from the University of Washington:
- Aylin Caliskan, associate professor in the Information School;
- Chirag Shah, professor in the Information School;
- Franziska Roesner, professor in the Paul G. Allen School of Computer Science & Engineering;
- Noah A. Smith, professor in the Allen School and the UW’s vice provost for AI;
- and Ryan Calo, professor in the Information School and the School of Law.
How alarming do you find the hacks announced by OpenAI and Anthropic?
Franziska Roesner: I do find them somewhat alarming — not due to the hypothetical risks from an anthropomorphized runaway AI, but because complex interconnected systems are being built and seemingly run without much in the way of standard safeguards and auditing. The resulting outcomes are unsurprising to security experts, but are sensationalized as AI risk.
Ryan Calo: The timing makes me a little skeptical. Is OpenAI trying to match Anthropic by arguing that its systems are just as scary? Is Hugging Face trying to look relevant in advance of its purchase by Nvidia? But yes — this sort of emergent behavior is concerning.
Noah A. Smith: We’ve been told that the beast got out of the cage, but we don’t know enough about the cage the beast was in. The demonstrations may establish an important new capability in these AI models without establishing the broader risk people are inferring. Assessing the underlying risk depends on what access, scaffolding, permissions and safeguards the system had. Anthropic’s own subsequent analysis shows that the alarming behavior depended heavily on what tools the model was given, what it was allowed to access, and how the experiment was set up, not just on the model itself.
Chirag Shah: I’m in half-agreement with scholars like Timnit Gebru who warn that the big AI labs are creating this scare to distract us from real problems that AI is causing. I also concur with Geoffrey Hinton and others who have been warning us about the security threats posed by the frontier models. I don’t think these two viewpoints are mutually exclusive: Yes, there are many other potential harms being created by AI, but the hacks and other security issues are real too and could be more devastating. Worse, we may not have time or opportunity to react, fix or reverse.
Aylin Caliskan: When such a complex system is equipped with tools and capabilities that enable it to interact with other complex systems, we should expect unforeseen exploits, problems and unintended consequences by default. The safety of these systems needs to be rigorously evaluated under controlled conditions and in real time, and appropriate guardrails should be dynamically integrated while they’re running.
What do you make of former Anthropic researcher Jacob Coxon’s claim that, “The people building AI earnestly believe that it could kill us all by the end of the decade”?
RC: I worry engineers like Mr. Coxon are playing into an industry rhetoric that would have society focus on speculative, existential threats, rather than immediate, real-world harms. I argued as much in 2023 in this magazine essay.
NS: I think most people don’t want to kill others or die themselves. Is he claiming that AI builders, collectively, want to harm others? Why are they building AI? Extraordinary claims about what AI builders collectively believe need evidence.
CS: I don’t buy it. I’d put this in the same category as the Y2K bug or communism destroying the world. AI has real benefits and dangers, but world-saving or world-destroying characterizations are neither realistic nor helpful.
AC: What does “believe” mean in Coxon’s sentence? Does it mean being unable to rule out a risk with 100% certainty, or does it mean that a large group of people building AI strongly believe that AI will be a net negative, yet continue to dedicate their resources to AI development? In theory, many things are possible. In practice, how likely are they?
FR: I wonder if these statements say more about the people making them than about the fundamental capabilities of AI. This article from science fiction writer Ted Chiang gives one perspective on this — that this belief in rampant, destructive AI is a product of the “no-holds-barred capitalism” practiced by major tech companies. It’s from 2017, but remarkably relevant.
Related
Sources for further reading, suggested by Noah A. Smith:
-
- Is AI actually going to kill us all? | WIRED
- The thrill of wanting AI to destroy the world | The Atlantic
- OpenAI and Anthropic unite against China’s open models | Axios
- An ex-Anthropic researcher claims AI could kill us all by 2030 — but he fails to answer the most essential question: What are we supposed to do about it? | Fortune
The people making these claims and announcements largely have financial stakes in these companies, which are both expected to IPO soon. How are you thinking about ulterior motives here?
CS: I see this as an attempt to steer the public into believing these companies are building world-changing tech that everyone needs to invest in or they’d miss out; that this tech would be so powerful that they rise up to national security level and gain power; and that the same tech could also be so dangerous that only they have the ability to curb it and they can self-regulate.
NS: It doesn’t take a conspiracy theorist to note that there are incentives at work. The financial stakes around prospective IPOs are enormous, and there are also long-standing concerns that safety arguments can shape regulation in ways that favor incumbent firms. Rules could reduce competition and independent scrutiny, concentrating both technological power and the authority to define what counts as “safe” in the hands of a few companies. They could also bar many people from participating in what the technology is designed to do, for example, by slowing or stopping work on open-source alternatives.
What should be done about AI risk?
NS: Risks need to be defined based on independent scrutiny and high-quality evidence, not messaging from organizations and people with a stake in what the response to risk looks like. We need sensible liability and accountability for harms, and governance proportional to demonstrated risks in real-world contexts rather than speculative narratives and science fiction. We should be especially wary of rules that entrench incumbent interests or treat closed, centralized control as synonymous with safety.
Openness is part of safety: If outsiders cannot inspect, reproduce and challenge claims about dangerous behavior, we are left trusting the organizations that have the strongest incentives to frame the narrative.
RC: Some combination of common law liability and regulation needs to create adequate incentives for AI companies to address the inevitable harms of this trillion-dollar industry.
FR: To me, the bigger question for safety is less, “What can AI models do in isolation?” and more, “How and why are we building these models into increasingly complex systems?” Computer systems security, for example, has already offered us examples of how to build these systems. More generally, we should all — whether we are building, integrating or using AI — anticipate how systems might be misused by people or harm them and adjust our systems accordingly.
AC: Academic freedom, independent evaluation and development, and open science play critical roles in analyzing and mitigating AI risks, as well as in effectively disseminating findings and evidence to inform policy and the public. To better manage risks, we should be designing AI deployment contexts in collaboration with stakeholders and communities, providing evidence to demonstrate net positive deployment effects that do not disproportionately benefit specific entities or groups, and iteratively identifying, isolating, and minimizing risks.
CS: Establish and fund commissions and taskforces that audit these companies and models and make independent assessments and recommendations. Make the companies rolling out these models accountable for any harms caused by their tech. Educate and empower the public through media, policies and democratic frameworks that give them a real say in what happens to their lives and labor through these technologies.
To set up an interview with an AI expert, contact Stefan Milne at stmilne@uw.edu.