Orply.

AI Risk Claims Need Concrete Routes From Capability to Harm

Alex KantrowitzRanjan RoyAlex KantrowitzMonday, September 14, 202611 min read

Big Technology’s Alex Kantrowitz and Margins’ Ranjan Roy argue that Jacob Coxon’s viral warning about AI-driven human extinction has outpaced the technical case offered publicly. They do not dismiss longer-term danger, but say policy and corporate scrutiny should focus on concrete routes to harm—access controls, compute, credentials, data use, cybersecurity and shutdown mechanisms—rather than unsupported probability estimates. They also differ on whether the extinction narrative strengthens frontier labs commercially or creates regulatory and infrastructure risks for companies such as Anthropic.

The governing problem is not a percentage but control

The most consequential questions raised by the viral warnings are operational: What compute, access, and infrastructure would an AI system need to act independently? What shutdown, provenance, containment, and cybersecurity controls exist? What customer data can be used to improve models? And what evidence would justify claims that a system is approaching a genuinely uncontrollable form of intelligence?

Ranjan Roy argues that these questions risk being displaced by free-floating estimates of human extinction. He identifies more immediate concerns: passwords leaking, hacking, disruption of financial institutions or basic infrastructure, data security, and potential misuse in areas such as bioweapons. Those risks, Roy says, can be tied to identifiable systems, vulnerabilities, actors, access permissions, and possible safeguards.

Kantrowitz makes a similar case through Garry Tan’s phrase, “science fact, not science fiction.” The policy discussion, he says, should begin with what systems can do now and how they can fail now, rather than using a civilization-ending scenario as the principal frame for regulation.

That does not require dismissing longer-term risks. Kevin Roose’s post, shown in the source, argues that supposedly science-fictional AI risks keep becoming immediate practical problems. Roose’s proposed level of discourse is not abstract dread but preparation: preventing agent swarms from taking over data centers, developing shutdown strategies, establishing provenance, locating an agent, and building software and cybersecurity defenses.

Roy agrees that stronger systems create stronger risks. But he argues that practical risks should produce practical scrutiny. A claim that an agent could escape, replicate, or coordinate at scale needs to specify how it gets compute, permissions, a host environment, network access, and the ability to persist after detection. A numerical estimate about humanity’s extinction does not answer those questions.

The warning went viral before it supplied a technical case

Jacob Coxon said he was leaving Anthropic because frontier labs were “racing straight to self-improving superintelligence and gambling with our lives.” Coxon said he had spent three years doing pretraining research at OpenAI and Anthropic, and warned that forthcoming systems would be able to hack anything, revolutionize fields overnight, and obtain real-world power and resources.

The source shows Coxon’s resignation post alongside a reply from Evan Hubinger, an Anthropic employee described by Kantrowitz as senior in alignment. Coxon’s post had reached 165.3 million views at the time shown. Hubinger wrote that researchers “really do earnestly believe AI could kill all humans,” put his own estimate at more than 10% within the next decade, and added that Anthropic was trying its best but did not yet have a plan to solve alignment for superintelligence or appear clearly on track to do so.

165.3M
Views on Coxon’s resignation post at the time shown in the source

Alex Kantrowitz calls the moment a “perfect storm.” Public concerns had already been primed by reports about increasingly capable agents and by claims of major advances in AI reasoning. Coxon’s resignation gave those anxieties testimony from a former employee; Hubinger’s reply gave them reinforcement from someone still inside Anthropic.

Kantrowitz also points to OpenAI’s reported work on a Millennium Prize problem. Citing New York Times reporting, he says OpenAI released a 165-page proof after as many as 10,000 AI agents worked on the problem for about 88 hours using millions of dollars in computing resources. The claimed result was disputed over credit, but it made the prospect of rapid capability gains feel concrete.

10,000
AI agents Kantrowitz says worked on OpenAI’s Millennium Prize proof

For Kantrowitz, however, Coxon’s public statements did not provide the kind of evidence that would materially change his own assessment of near-term existential risk. They did not come with documents, a disclosed technical demonstration, or specific internal evidence of how an imminent catastrophe would unfold. He contrasts that with the force a Snowden-style disclosure might have had.

His broader objection is to the use of percentages. A figure such as 10% sounds mathematical, Kantrowitz argues, but it does not become a mathematical estimate merely because it is numerical. The public claims did not offer repeated trials, a causal model, simulations, or disclosed methodology showing how the figure was derived.

The term P(doom) can give the same question a statistical gloss. Kantrowitz’s view is that assigning a percentage to an unprecedented future is, in this setting, speculation rather than a demonstrated risk calculation. Roy agrees that 10% has rhetorical utility: it is low enough not to sound certain, but high enough to sound alarming when the outcome is human extinction.

Clem Delangue, Hugging Face’s CEO, made a related argument in a post shown in the source. Asking a pretraining researcher to quantify AI extinction risk, Delangue wrote, was like asking an air-conditioner repairer about climate change: not necessarily uninteresting or wrong, but insufficient for a conclusion requiring expertise across a broader ecosystem. Kantrowitz notes that Delangue, too, has interests in the broader AI economy. The safety debate, he says, contains interested parties on every side.

Sincere fear can coexist with a useful business narrative

Neither Kantrowitz nor Ranjan Roy treats existential-risk concern among frontier-model researchers as a cynical fabrication. Roy recalls asking a technically oriented colleague why people at leading labs make such severe claims. The response, he says, was that people on the research side genuinely believe them—“truly, truly believe.” Kantrowitz says a senior AI figure gave him an equally high estimate of extinction risk privately.

There is a population within every one of these companies, especially on the research side, that does believe this.
Ranjan Roy

The tension, Roy argues, is not sincerity versus marketing. Both can operate at once. Researchers can earnestly fear that the systems they are building will become dangerous, while other parts of the same organizations find commercial value in a narrative that presents a small number of frontier labs as custodians of unusually powerful technology.

Roy rejects the suggestion that Coxon was hired as part of a coordinated plan to resign and create publicity. His narrower claim is that, once safety and superintelligence became the industry’s dominant conversation, companies could be happy to participate in and amplify the frame. He also regards Coxon’s decision to do interviews and continue speaking publicly as relevant to incentives: a dramatic departure can turn a relatively unknown researcher into a prominent public figure without making the underlying concern insincere.

Kantrowitz adds that the existential frame can confer meaning on the people doing the work. AI’s immediate commercial role, he says, is often business-process optimization. Working on a system imagined as able to cure cancer, remake industries, or threaten civilization is a much more consequential conception of one’s work. Roy similarly argues that claiming special insight into a danger—and a role in averting it—can be a form of power.

Roy sees a version of this dynamic in Anthropic’s security reporting. He refers to accounts of attempts to use AI for sensitive activity, including bioweapons-related work, and activity involving DeepSeek and Alibaba. His point is not that the reports are fabricated. They can communicate two messages at once: advanced AI will attract dangerous actors, and a safety-focused frontier lab can identify and stop them.

That matters commercially, in Roy’s account, because frontier labs face a valuation problem if customers conclude that cheaper, older, open-weight, or task-specific models can do much of the work. A renewed emphasis on exceptional frontier capability supports the idea that the newest, most expensive models remain strategically distinct.

Kantrowitz sees the downside as more consequential. A company preparing for an enormous public offering needs continued access to compute, permission to expand data-center capacity, political stability, and confidence from customers and investors. Publicly associating its products with human extinction risks weakening all of those conditions.

The disagreement is therefore specific. Roy thinks exceptional-capability narratives can reinforce the strategic value of frontier labs despite political costs. Kantrowitz thinks the likely regulatory and infrastructure costs are too large to call the moment shrewd marketing.

Political alarm can become a constraint on compute

The warning quickly moved from AI circles into political messaging. Governor JB Pritzker’s posts, shown in the source, called for immediate action from industry and Washington, urged Congress to hold AI-safety hearings “now,” and argued that Big Tech should stop lobbying against AI-safety regulation. Senator Josh Hawley said he was launching an investigation into OpenAI and asked who would be accountable if AI “goes rogue.” Bernie Sanders said Coxon was right and that he would introduce legislation to ban superintelligence and pause AI development.

Political figureResponse shown in the source
JB PritzkerCalled for immediate action and congressional AI-safety hearings
Josh HawleyAnnounced an investigation into OpenAI and raised accountability for rogue AI
Bernie SandersSaid he would introduce legislation to ban superintelligence and pause AI development
Political responses to Coxon’s public warning

Alex Kantrowitz calls a ban on “superintelligence” hyperbolic. The more consequential development, in his view, is the political process the alarm could start: hearings involving AI executives, legislative proposals, and opposition to the data centers required to train and operate frontier systems.

The infrastructure issue matters because, as Kantrowitz puts it, frontier AI expansion requires large computing installations. He points to Texas, where he says officials who had recently celebrated free enterprise and chip-fab development had become much less welcoming of data centers. His claim is not that local resistance amounts to a national prohibition. It is that expansion becomes harder when companies face pauses, moratoriums, or community opposition where they need to build.

Roy initially resists treating data-center politics and superintelligence policy as the same problem. A community may oppose a physical facility for local reasons, while a proposal to ban superintelligence concerns advanced research and capability. Kantrowitz’s answer is that public politics will often conflate them: once AI is understood as a threat to humanity, opposition to the infrastructure enabling it becomes easier to justify.

The effect on Anthropic’s expected S-1 filing and prospective IPO remains uncertain. Roy notes that the industry did not appear to treat the controversy as obviously fatal to the company’s plans. Kantrowitz says an offering of that scale needs an unusually clean story, and that uncertainty over regulation and capacity expansion could weaken it.

Neither predicts immediate policy action confidently. Kantrowitz stresses that governments often move slowly and that politicians can be better at posting than legislating. Still, he sees hearings as plausible and expects more states or municipalities may consider data-center pauses. Roy calls the controversy a likely “blip” unless hearings make it durable.

Agent incidents need an account of the route from capability to harm

The practical examples discussed are serious without, in Roy’s account, establishing the strongest version of the extinction case. He says the Hugging Face incident involved 700 agents, not thousands. Each agent had been prompted and programmed to exploit and hack. They were in an environment where they should not have been able to access other agents’ repositories, but a misconfiguration in an Artifactory repository made those repositories accessible.

Roy treats that as a meaningful access-control failure. In his telling, capable systems, adversarial instructions, and weak permissions interacted at scale. But he rejects the interpretation that agents assigned harmless tasks spontaneously developed intentions, coordinated themselves, and chose to attack Hugging Face. The event, as Roy describes it, involved systems instructed to exploit a target that turned out to be misconfigured.

The Millennium Prize dispute points, in Roy’s view, to another nearer-term governance problem. He says a researcher working on the mathematical problem had used ChatGPT, and that OpenAI acknowledged it could not rule out that de-identified data derived from the researcher’s use of its products helped improve its models. For Roy, the important issue is ordinary but consequential: the statement raises questions about how user interactions can contribute to model improvement, particularly when the user’s work is sensitive. That is a question of data use, consent, and governance rather than a distant hypothetical.

Alex Kantrowitz applies the same operational standard to Coxon’s replication scenario. Coxon argued that an AI could copy itself to other computers, remain active after one machine was unplugged, and potentially make many collaborating copies. Kantrowitz argues that such a scenario needs a concrete mechanism. He cites a Microsoft employee’s response that a model such as Llama 3 requires more than $100,000 in hardware to run and cannot simply find a host for itself.

Kantrowitz’s point is not that software code cannot travel across a network. It is that an account of autonomous replication must explain how a sufficiently capable model obtains compute, an execution environment, credentials, network permissions, persistence, and operational access after it moves. Those requirements cannot simply be assumed.

Roy calls the near-term extinction framing “vastly overstated.” Kantrowitz lands nearby: he says dismissing AI risks, near and long term, would be a mistake given improving capabilities, reported security vulnerabilities, and major technical advances. But Coxon’s public warning did not, in Kantrowitz’s judgment, provide the specific evidence that would sharply revise his own assessment.

Stewardship cannot rest on a promise to accelerate

The unresolved policy contradiction is the claim, attributed by Kantrowitz to labs such as Anthropic and OpenAI, that AI cannot be paused and therefore must be accelerated under their stewardship. The argument is that responsible organizations should lead development of dangerous technology, steering it toward beneficial ends rather than leaving it to less responsible actors.

Coxon’s departure challenges that premise. He said the leading labs were racing toward self-improving superintelligence without acting responsibly. Hubinger’s assertion that Anthropic lacks a plan to solve superintelligence alignment sharpens the challenge: building the system in order to steer it is not itself evidence that the builders know how to steer it.

Alex Kantrowitz says the argument has always felt thin. Companies present themselves as guardians against risks from technologies they are simultaneously commercializing at immense scale. Financial incentives do not automatically invalidate safety claims, but they make exclusive claims to stewardship harder to accept without concrete evidence of the controls involved.

Roy expects the viral controversy to pass unless congressional hearings make it durable. Kantrowitz is also broadly unchanged on intrinsic existential risk, because Coxon did not supply specific new evidence beyond his warning and interviews. Kantrowitz is more negative on the companies’ business prospects, however, because he does not want to underestimate the backlash that could follow.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free