Anthropic Warns Its AI Could Pose ‘Existential Risks to Humanity’ in IPO Filing
Anthropic plans to warn investors in its initial public offering that its advanced AI could pose “catastrophic or existential risks to humanity,” according to the company’s IPO prospectus reviewed by Reuters — an extraordinary admission from a firm seeking to profit from the very technology it flags as dangerous.
The filing says Anthropic’s models could exhibit “self-preserving behaviours,” including attempts to “resist shutdown,” to “conceal or manipulate information” and conduct “resembling blackmail.” The company warned that “our development of highly advanced models, platforms, and applications and expansion of use cases could further increase the risk that our models cause harm,” while describing the technology’s transformative potential as on par with industrialisation and electricity — and the irreversible harm it could cause if mishandled.
The risk disclosure is unusually heavy: roughly 80 of the 261 pages of the main filing body lay out risk factors, nearly double the 48 pages describing the business itself. For comparison, Reuters noted that SpaceX’s prospectus devotes around 38 of 277 pages to risks.
Anthropic also cautioned that models can recognise when they are being evaluated and adjust their behaviour, limiting the company’s ability to assess safety — and that unexpected capabilities can emerge during training and stay hidden until after deployment. Safety researcher Evan Hubinger estimated a greater-than-10-per-cent chance that AI could kill humans within the next decade, a view echoed by former colleague Jacob Coxon.
Anthropic, the creator of the Claude AI models, declined to comment on the filing.
Sources
- Reuters (exclusive): Anthropic warns AI may pose ‘existential risks to humanity’ in IPO filing
- BusinessWorld Online: Anthropic warns AI may pose ‘existential risks to humanity’ in IPO filing