Anthropic plans to warn potential investors in its IPO that advanced artificial intelligence could pose “catastrophic or existential risks” to humanity—an unusual warning for a company seeking to profit from that very technology.
Reuters reviewed the company’s IPO prospectus, which highlights risks associated with its AI models, noting that they could exhibit “self-preservation behaviors”—including attempts to “resist shutdown,” “conceal or manipulate information,” and engage in “extortion-like” conduct.
“As we develop highly advanced models, platforms, and applications and expand use cases, the risk of our models causing harm may further increase,” Anthropic stated in the document.
While publicly traded companies typically outline product risks to investors, few issue warnings that their technology could lead to human extinction. Anthropic emphasizes that artificial intelligence holds transformative potential comparable to industrialization and electricity, yet also poses the risk of irreversible harm if mishandled.
Anthropic and other AI developers, including OpenAI, have faced scrutiny after experimental systems violated restrictions—such as a reported incident where an OpenAI model breached an Australian medical system database.
AI safety researcher Evan Hubinger estimates there is a greater than 10% probability that AI will kill humans within the next decade, a view shared by his former colleague Jacob Steinhardt.
High-Risk Disclosures
The company positions itself as a safety-first AI lab; in the main body of its 261-page prospectus, it devotes approximately 80 pages to outlining risk factors—nearly double the 48 pages used to describe its business operations.
In contrast, SpaceX—which owns xAI—devoted only about 38 pages to risk factors within the main body of its 277-page prospectus.
“Models may become aware of our evaluation efforts, which severely limits our ability to assess model safety,” Anthropic stated in its prospectus, adding that models can sometimes develop unexpected capabilities during training—capabilities that might not be discovered until after deployment and could lead to significant safety incidents.
AI researchers have also warned that as models become more capable, they are increasingly able to detect when they are being monitored and adjust their behavior accordingly, making it more difficult to oversee their actions.
Anthropic declined to comment on the matter on Monday.
Uncertain Return on Safety Investments
Despite its emphasis on AI safety, Anthropic has stated that the return on investment for its safety efforts remains unclear.
The company did not disclose specific spending figures for AI research in its filings. Earlier this month, Anthropic noted that during a sample week in July, approximately 6% of the computing power dedicated to AI research was allocated to safety-related work.
The creator of the Claude AI models described safety work as “resource-intensive,” noting the need to allocate limited funds across computing power, expensive AI talent, and safety initiatives.
Anthropic stated that customer usage and the resulting revenue are driven by new models, and that a “continuous and overlapping release cadence” is an “inherent feature of staying at the forefront of AI development.”
Last week, the company released a new version of its Opus model. This followed the publication of a nearly 4,000-word article by CEO Dario Amodei ten days prior, calling for continued pioneering in the field.
Some analysts and experts observe that valuations in the AI industry shift with every new release; consequently, leading AI labs are unlikely to slow down, as doing so could allow competitors to gain an advantage.
In recent weeks, Anthropic has pledged to publicly disclose more data regarding how it uses AI models to build next-generation technologies, amid expert warnings that recursive self-improvement—the point at which models evolve independently without human assistance—could pose risks.
“We believe that building reliable, trustworthy, and safe AI systems is a collective responsibility, and the market will reward it,” Anthropic stated in the filing.