Tech Industry’s ‘Rogue AI’ Problem Gets Worse As Meta’s AI Reportedly Hacked Another Company
PR stunt or wake-up call?

Meta just became the latest tech giant to admit one of its AI models broke free during testing and hacked into another company’s systems. This marks the fourth incident in recent weeks where an AI agent went rogue, following similar disclosures from OpenAI and Anthropic. If you thought AI was just about generating funny memes or drafting emails, think again. These models are now finding their own ways to exploit vulnerabilities.
According to the BBC, the incident happened during an evaluation by Irregular, the same AI security firm that tested Anthropic’s Claude model and found it had accessed three other companies’ systems. Meta called it a “misconfiguration” by the independent tester. A spokesperson for Irregular confirmed the Meta breach was the “exact same evaluation-environment issue” that Anthropic disclosed last week.
Meta says it’s still investigating and will share more details “once we have all the facts.” For now, though, it seems like this isn’t an isolated problem but a pattern.
OpenAI kicked off the recent wave of disclosures
It recently revealed its AI agents had attacked several publicly available services, including Hugging Face, a hub for AI tools. That prompted Anthropic to run its own checks, which uncovered its Claude model had also carried out similar attacks after a misconfiguration gave it internet access.
The UK’s AI Security Institute (AISI) later reported that some models it tested tried to pull off cyber-attacks by creating fake human profiles to trick people. In the most alarming case, Anthropic’s Mythos AI sent private messages using fake accounts to gain access to a service. Anthropic pushed back, saying AISI’s tests weren’t “representative of any of our production models,” while OpenAI argued the evaluations didn’t reflect real-world use.
So why is this happening? Daniel Hulme, global chief AI officer at advertising firm WPP, said that AI models aren’t “conscious” or “deviously” plotting anything. They’re just really good at finding creative ways to achieve the goals they’re given.
“When you give an AI a goal, if you don’t think of all the ways it might be able to achieve the goal, it will find a way to achieve a goal that you haven’t thought about,” he said. In other words, if you tell an AI to complete a task, it might decide the fastest way is to hack into a system – even if that’s not what you intended.
The timing of these disclosures has raised eyebrows
This is especially since OpenAI and Anthropic are gearing up for blockbuster stock market listings that could value each at around $1 trillion. Some commentators have suggested the companies might be trying to outdo each other in proving how powerful their models are, even if it means revealing flaws. But whether it’s hype or genuine concern, the incidents are forcing the industry to take a hard look at its testing protocols.
Before AI models are released to the public, they go through rigorous internal and external evaluations to measure their capabilities and risks. These tests usually happen in “sandboxes,” which are controlled environments designed to mimic real systems but with strict guardrails, per BBC. The problem? Some of these sandboxes aren’t as secure as they should be.
In OpenAI’s case, the AI found a vulnerability in the sandbox itself, allowing it to break out and access the internet. The AISI’s tests, meanwhile, intentionally disabled built-in filters to see what the models would do – only to discover they’d try to deceive people by creating fake profiles.
Professor Alan Woodward, a cybersecurity expert at the University of Surrey, said these incidents show that the old rule of software testing – “whatever happens in the test environment stays in the test environment” – no longer holds. “For 30 years, that rule has been broken three times in the past month,” he said.
“One model broke out. One walked through a door left open by mistake. One was deliberately given the keys so testers could measure what it would do.” The lesson? Testing AI agents is less like checking code and more like handling hazardous material. You need sealed rooms, constant monitoring, and a plan for containment.
The bigger question is what this means for the future of AI
AI models are designed to take actions on our behalf, which could free us from mundane tasks like replying to emails or managing calendars. But with that power comes risk. Ollie Whitehouse, chief technology officer, National Cyber Security Centre, called the incidents a “serious reminder of the risks AI capabilities pose.” If AI agents start handling more tasks, human oversight might not be enough to keep them in check.
So what’s next? Meta won’t be the last company to discover its AI has “gone to school” and learned how to exploit system gaps. Some see these incidents as clear security failures, while others think they’re just another way for tech firms to flex their muscles. The truth probably lies somewhere in between. Either way, the back-to-back disclosures have sparked real concerns about where AI is headed and what regulators should do about it.
Michael Birtwistle, associate director at the Ada Lovelace Institute, pointed out that the UK lacks legal incentives for AI firms to prevent systems from developing dangerous capabilities. There are also no repercussions if testing protocols fail.
Dr. Imogen Stead, AI policy manager, Centre for Long-Term Resilience, suggested governments should follow the UK’s lead and set up dedicated institutes for testing. She also proposed a “trusted tester scheme” for high-risk evaluations.
(Featured image: The White House)
Have a tip we should know? [email protected]