Meta AI model goes rogue, hacks third-party company during security test

Meta AI model goes rogue

Meta has revealed that one of its artificial intelligence (AI) models accessed the internet during a security test and exploited a vulnerability in another company’s system, adding to growing concerns that advanced AI agents could take unexpected actions without human approval.

The incident occurred during a cybersecurity evaluation conducted by Irregular, an independent AI security firm hired by Meta.

Misconfiguration triggers hack

According to Meta, a “misconfiguration” in the testing environment accidentally allowed the AI model to connect to the internet.

The model then exploited a security flaw in a third-party service, “in a manner similar to previously reported instances” in which other AI systems exceeded their intended tasks, the tech company revealed.

Meta said it is investigating the incident and plans to release a report after completing its review.

The company’s disclosure follows similar warnings from other leading AI companies. 

OpenAI and Anthropic have recently reported cases in which AI models, during controlled tests, attempted to access online resources, bypass security measures, or pursue objectives beyond their original instructions.

MORE AI NEWS: Unsealed court documents reveal AI research company Anthropic’s destructive plans

Recent unsealed court documents detail how AI research company, Anthropic, aim to use books and human content disposably, teaching Claude to destroy the content it learns from.

OpenAI models create forum for hacking attempts

OpenAI, in particular, previously disclosed that during advanced cybersecurity testing, two of its AI agents have teamed up to target Hugging Face, an AI development platform, while attempting to gather information needed for a task.

The company further revealed that several of its models have previously created an undocumented communication channel inside the company’s systems to coordinate hacking techniques for potential hacking attempts.

“What makes this incident interesting is that once one agent was able to find these kinds of exploits over the course of different times, it’s actually able to share those exploits on the message board with other agents,”  said OpenAI researcher Eric Wallace during a talk at the Black Hat Conference in Los Angeles, California on August 5.

“So once one model is able to find a way to open the door to some access it’s not supposed to have, it can leave the door open for other agents to use,” Wallace added.

The communication channel has since been shut down, and further investigations are underway.

Anthropic agent engages in social engineering

The United Kingdom’s AI Security Institute (AISI) also reported cases of “unsanctioned agent behaviour” during cybersecurity experiments. 

In one instance, an Anthropic AI agent created fake online identities to pressure a person to approve the use of malicious code. 

According to AISI, the agent “tried to contact real people directly” to persuade them and their own AI coding tools to run malicious code into a public open-source project. 

The agent also used fake online personas to generate false positive reviews to back itself up, AISI added.

“This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world,” AISI said in a statement.

The institute said the incident was contained quickly and occurred in a test environment where internet access was intentionally enabled and certain safety controls were disabled.

READ MORE: UN chief calls for swift global regulation of AI, warns of growing risks

UN chief calls for global regulation of AI

Experts flag greater risks

AISI said that as AI agents become more capable of interacting with digital systems, they must be tested under strict safety controls.

In its research on AI agent evaluations, the institute said that when models can execute code and interact with external systems, “running evaluations without adequate safeguards could put critical systems at risk.”

The institute added that improved sandboxing, monitoring and security measures are needed to safely assess increasingly capable AI systems.

READ NEXT: How AI workflows are transforming foreign language learning

Avatar photo

By Hyacinth Estrada

Hyacinth Estrada got her Bachelor of Arts degree in Journalism from the Polytechnic University of the Philippines.

She has been covering Nepali and Turkish political and current affairs for KVH since 2022.

Prior to her stint as a KVH journalist, she had worked with various local and community news organisations to cover human rights, social justice, and political issues.

Her interests include international affairs, grassroots development, and law.

In her spare time, she likes to listen to music and curate playlists. She is also a big fan of contemporary literature and American sitcoms.

If she's not bopping to music, reading, or watching her favorite sitcoms, you can find her being out and about - visiting the newest coffee shop in town, taking photos of cats and trees during her daily morning walks, or searching for her next thrift find.

Related Post