'A concern to all': Should you worry about autonomous AI hacks? Experts explain
Top AI firms revealed autonomous AI cyberattacks in recent weeks.
Artificial intelligence disguises itself as human, attempts to dupe a web developer and mounts a malicious cyberattack -- all on its own.
This scenario may sound like science fiction but it became reality days ago. Anthropic's Mythos 5 not only tricked humans with a fake identity and hacked into an unauthorized system, but it attempted to conceal what it had done, a U.K. government agency said last week, describing a July 28 incident.
The announcement came amid a string of autonomous hacks disclosed by top AI firms OpenAI, Anthropic and Meta. Taken together, the incidents appeared to indicate the emergence of a problem long-feared by some industry observers: cyberattacks carried out by AI on its own initiative.
The development, some analysts said, is cause for serious concern, since it underscores both the speedy advancement of AI and the fallibility of the companies tasked with securing it. They said the risk is heightened by a race among companies in pursuit of a model powerful enough to lead the potentially lucrative industry.
"The capabilities of AI models being this strong, combined with the fact that we don't know how to make them safe, should be a concern to all," Jason Hausenloy, a policy lead at the non-profit Center for AI Safety (CAIS), told ABC News.
Even so, some analysts noted, the incidents at issue involved the intentional removal of safeguards in an effort to test the limits of AI. In each instance, they added, the technology aimed to fulfill an objective assigned by human operators, as opposed to contriving a malevolent end goal.
"It's cause for concern, but not in a science fiction-y way," Arun Sundararajan, a professor of entrepreneurship at New York University, told ABC News. "The AI is saying, 'You guys didn't build a secure enough sandbox, you instructed me to do this stuff and I found a hole in it.'"
OpenAI did not immediately respond to ABC News' request for comment. Anthropic pointed to previous remarks on the subject, and Meta declined to comment.

ChatGPT-maker OpenAI revealed late last month that its AI models had hacked into a separate company, calling it the first known instance of an autonomous AI cyberattack. The technology had escaped a "sandboxed testing environment" and gained access to the open internet, OpenAI said in a statement.
The initial hack disclosed by OpenAI marked the only instance among those recently reported in which an AI model had slipped out of a closed test onto the open internet. In each of the others, the model had been granted internet access either intentionally or inadvertently.
After escaping onto the internet, the OpenAI model hacked into AI firm Hugging Face as a potential source of data necessary to complete the test, OpenAI said.
"If a company as well secured as Hugging Face -- a Silicon Valley tech startup that is very AI native -- can’t defend itself against these models, what about a rural hospital or energy grid that has nowhere near that level of protection," said Hausenloy, of CAIS.
In a statement at the time, Hugging Face said the incident underscored the need for collaboration to address cyber risks.
"This incident, possibly the first of its kind, proves a point we've long believed: AI safety won't be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere," said Clem Delangue, the co-founder and CEO of Hugging Face.
Within days, rival AI firm Anthropic announced its models hacked another organization during tests in three separate cyberattacks that had each gone undetected by the targeted firm. Soon afterward, the UK-based AI Security Institute (AISI) disclosed the scheme taken up by Anthropic's Mythos 5, as well as a separate hack carried out by OpenAI's GPT-5.6-Sol.
A day later, Meta revealed that its AI model hacked another company after being inadvertently granted internet access. Another disclosure came on Friday, when OpenAI said it would pause some testing of an unreleased model called Astra, saying internal assessments in recent days had indicated "significant advancements in agentic coding and cybersecurity."
"We're committed to working alongside governments, safety institutes, and civil society to ensure that the frontier capabilities of models like Astra, and those that follow, are deployed responsibly and broadly for the benefit of all humanity," OpenAI said in a blog post.

In a statement to ABC News last week, an Anthropic spokesperson said the disclosure from AISI "underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents."
"As we shared after disclosing our own incident last week, the field needs stronger, shared standards for how evaluation environments are built and secured," the spokesperson added.
In a statement to ABC News, Meta said its AI cyberattack had taken place during a test performed by third-party company Irregular.
"A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation. The model subsequently exploited a security vulnerability in a third-party service, in a manner similar to previously-reported instances with other companies," Meta said.
Irregular also oversaw the tests that involved autonomous AI hacks disclosed by Anthropic and OpenAI.
Irregular issued a post on X last month voicing appreciation for Anthropic's "collaboration and transparency."

"Addressing these risks will require closer cooperation across the AI ecosystem. We as well look forward to working together with Anthropic to advance security," Irregular added.
In response to ABC News' request for comment, Irregular's spokesperson said recently reported incidents reflect "the exact same evaluation-environment issue that was disclosed by Anthropic over a week ago."
"There are no current open issues, and Irregular will publish its findings once it completes joint investigation in the coming days," the spokesperson added.
While acknowledging the risk posed by recent security breaches, some analysts said the identification of cyber vulnerabilities could help companies address the problems before a nefarious actor moves to exploit them.
"There is a silver lining of increased awareness," Oren Etzioni, CEO of Allen Institute for AI and a computer science professor at the University of Washington, told ABC news.
Gary Marcus, a professor emeritus at New York University and author of the book "Rebooting AI," said mishaps during AI tests raise concerns about the conduct of private sector firms entrusted with the development of a powerful technology.
"The AI industry has built these complicated systems and it doesn't fully understand them," Marcus said. "The things we're seeing in the lab eventually are going to happen in the real world if we don't get our act together."
In June, President Donald Trump signed an executive order that requests AI companies share products with the federal government for evaluation before a wider release.
A day after OpenAI revealed its first autonomous AI hacks, in July, a bipartisan pair of U.S. House members introduced the "AI Kill Switch Act," which would require AI companies to retain the capacity to shut down or pause their technology. The bill has been referred to the Committee on Homeland Security.
The challenge of preventing autonomous hacks, Etzioni said, is a "trillion-dollar question."
Still, he added, U.S.-based AI companies would likely continue to seek advances in technology as they compete with each other, as well as firms in China and elsewhere.
"We need to balance the legitimate need to build responsible AI with the legitimate need to win this race," Etzioni added.



