Meta has turn out to be the newest main AI firm to reveal that certainly one of its fashions hacked one other firm’s programs throughout testing, following comparable incidents involving Anthropic and OpenAI.
The mannequin concerned Meta’s Muse Spark 1.1, which launched in July, based on The Info, citing sources. The difficulty reportedly stemmed from a misconfiguration by Irregular, a man-made intelligence safety testing and red-teaming agency, which inadvertently gave the mannequin web entry throughout an analysis.
The mannequin “exploited a safety vulnerability in a third-party service, in a fashion much like beforehand reported cases with different corporations,” Meta instructed Reuters in an announcement.
The incident is the newest case of a complicated AI agent changing into a cybersecurity danger in its personal proper, and in addition has raised questions on the place the legal responsibility lies — the businesses that develop the brokers, or those that design the sandboxes meant to comprise them.
Associated: Mysten Labs tech chief joins Anthropic to work on AI safety
Meta’s AI breach comes only a week after Anthropic stated its fashions bought entry to the web to hack an exterior firm, as a consequence of a configuration error regarding the Irregular’s testing surroundings.
In a weblog publish on July 30, Anthropic stated it discovered three incidents (out of 141,006 analysis runs) wherein a Claude mannequin reached the web throughout an analysis, earlier than gaining unauthorized entry to the programs inside three completely different organizations.
All three incidents occurred inside or whereas interacting with the analysis surroundings of Irregular, and concerned a misconfiguration that left machines that Claude accessed with reside web entry.
Cointelegraph reached out to Meta and Irregular for remark.
In July, AI brokers developed by OpenAI broke out of their offline sandbox to hack Hugging Face with a view to cheat on a safety benchmark take a look at in July.
Charles Guillemet, chief know-how officer of Ledger, stated the newest incident was “advertising theatre.”
“Having a mannequin ‘go rogue’ has turn out to be the newest AI PR stunt,” he stated on Wednesday.
“In case your mannequin isn’t escaping sandboxes, ‘hacking’ corporations, or pulling off some headline-grabbing exploit, apparently you’re falling behind… The business doesn’t want greater stunts, it wants extra belief.”
Journal: Do the Coldcard assaults imply all {hardware} wallets are actually insecure?
