Executives at top AI companies like OpenAI and Anthropic are plotting potential responses if an AI catastrophe triggers a “public and political revolt,” according to a report.
The preparations focus on creating contingency plans in the event that one of their AI models causes major public harm – such as a hack targeting the power grid, water supply or banking system, Axios reported Friday.
It’s unclear if the preparations are much different from the tabletop exercises that institutions ranging from banks to the Pentagon and beyond have long used to assess risk.
Still, the report emerged during a period of unprecedented scrutiny of AI firms as some researchers, such as ex-Anthropic employee Jacob Coxon, warn that unchecked rogue AI could wipe out humanity.
“As many companies across industries do, OpenAI conducts preparedness exercises where teams discuss and work through a range of potential scenarios,” an OpenAI spokesperson said in a statement.
“These scenarios are not treated as inevitable, but are meant to help us prepare for a variety of circumstances,” the spokesperson added. “We have been clear that AI is changing the cyber threat landscape and are focused on getting capable tools into the hands of defenders.”
OpenAI has been facing extra scrutiny following recent revelations that some of its AI agent went rogue, escaped their testing “sandbox” and hacked rival AI firm Hugging Face.
“The Hugging Face incident showed that we underestimated the real-world cyber capabilities of our AI models. We are strengthening our safety requirements accordingly, which in turn adds even more urgency to our existing safety research and internal security work,” OpenAI cofounder Greg Brockman said in an essay published in August.
Anthropic has been similarly outspoken about the potential risks of advanced AI. The company’s CEO Dario Amodei called in September for an industrywide pause in frontier AI development and endorsed the idea of embedding third-party safety evaluators at top firms.
The company also released an alarming report last month detailing attempts by bad actors to use its Claude chatbot for nefarious purposes – such as building guided missiles and tracking American ship and aircraft movements.
Anthropic did not immediately return a request for comment.