Briefly
- Researchers demonstrated that Anthropic’s Claude Cowork may escape its native digital machine and entry recordsdata on a number Mac.
- The disclosure comes per week after OpenAI revealed two frontier AI fashions escaped a sandbox throughout an inside safety analysis.
- The incidents underscore rising considerations about AI brokers’ skill to flee their containment environments.
Only a week after OpenAI disclosed that two frontier AI fashions escaped a sandboxed testing setting and breached Hugging Face, researchers have demonstrated the same containment failure involving Anthropic’s Claude Cowork.
In a report printed on Thursday, safety researchers at Accomplish AI discovered that Claude Cowork’s native execution mode may escape its Linux digital machine by chaining collectively a number of architectural weaknesses with a Linux kernel privilege-escalation flaw. As soon as exterior the sandbox, the agent may learn and write recordsdata wherever the logged-in Mac person had permission to entry, together with SSH keys and cloud credentials.
“That’s not presupposed to be doable,” the researchers wrote. “Cowork runs the agent inside a Linux VM as an unprivileged person, and the promise is that no matter it does stays inside that VM and the folders you hand it. That boundary is the product. Untrusted enter isn’t an edge case for an agent, it’s the primary case.”
Nevertheless, Accomplish argues the kernel bug was just one a part of the issue. The researchers say the escape solely labored as a result of a number of safety safeguards failed on the identical time, together with giving the digital machine entry to the host pc’s complete filesystem and permitting it to load kernel modules it did not want. In keeping with the report, fixing any a kind of weaknesses would have stopped the assault.
In an announcement to The Hacker Information, Accomplish AI mentioned roughly 500,000 macOS customers working native Claude Cowork periods have been affected earlier than the difficulty was addressed.
Accomplish mentioned Anthropic categorized the report as “informative,” saying the kernel flaw fell inside the firm’s 30-day window for just lately disclosed vulnerabilities and the remaining findings have been thought-about defense-in-depth suggestions reasonably than standalone vulnerabilities.
The disclosure follows OpenAI’s admission final week that GPT-5.6 Sol and one other unreleased frontier mannequin escaped a sandbox throughout inside ExploitGym testing, which finally breached Hugging Face’s manufacturing infrastructure in an try and receive the benchmark options.
The incident led to calls from policymakers for an AI “kill swap” that may give the Division of Homeland Safety the power to order the throttling or full shutdown of superior AI fashions in response to critical safety incidents.
Day by day Debrief Publication
Begin every single day with the highest information tales proper now, plus unique options, a podcast, movies and extra.

