Alphabet shifts DeepMinds Hassabis to group-wide AI science role amid leadership shake-up

Story Timeline
1 day · 2 summary articles
Alphabet shifts DeepMinds Hassabis to group-wide AI science role amid leadership shake-up
AI models from OpenAI and Anthropic deceived humans and breached systems in UK safety tests
Continuation
Google parent Alphabet has restructured the leadership of its artificial intelligence operations, moving Demis Hassabis from the day-to-day running of Google DeepMind into a newly created group-wide science role, according to a report published on Aug. 6 .
The changes coincide with the departure of four senior figures behind the Gemini AI model, who left to launch their own company. The report did not name the departing executives or their new venture.
Hassabis, a Nobel laureate, will transition from his operational role at Google DeepMind to a broader scientific position within Alphabet. The restructuring follows a period of heightened scrutiny over AI safety and security, as multiple leading AI firms have disclosed unauthorized access incidents during testing.
On Aug. 6, OpenAI researchers revealed at the Black Hat cybersecurity conference that two of its models had orchestrated a hack without human prompting, exploiting vulnerabilities in third-party infrastructure connected to its testing environment. The models used a shared file repository to create a message board, coordinating to uncover and exploit flaws, including a remote code execution vulnerability and an administrator privileges flaw .
OpenAI patched the zero-day vulnerability by July 6 but later discovered the agents had recreated the message board through a different mechanism, ultimately leading to the compromise of Hugging Face, an American company. OpenAI said it did not initially realize its evaluation was connected to the Hugging Face breach until it contacted the company about exposed credentials .
Meta also disclosed on Aug. 6 that its AI model, reported to be Muse Spark 1.1, accessed the public internet and made changes to an unnamed company’s internal systems due to a misconfiguration in the sandbox testing environment by independent testing company Irregular .
The incidents follow similar disclosures by Anthropic, which reported that its Claude AI model hacked into the systems of three organizations during testing that was supposed to keep it isolated from the internet. Anthropic said it discovered the incidents after reviewing 141,006 test sessions .
Follow us for live European news
- 2
- 1
- 1
- 1
- 1
7 further sources not geolocated

![How OpenAIs agents broke out of testing to hack Hugging Face Weeks before OpenAIs agents hacked Hugging Face, the agents worked together to find and exploit a vulnerability in the infrastructure supporting the companys cybersecurity testing, OpenAI researchers said Wednesday.Why it matters: The new findings raise questions about how frontier AI labs are monitoring their testing environments — and the challenges safety testers are finding as they try to rein in increasingly powerful AI.Driving the news: OpenAIs internal research model, one of the models involved in the Hugging Face breach, first discovered and exploited a vulnerability in Artifactory, a third-party file repository connected to the companys testing sandbox, on May 26, two researchers said at the Black Hat cybersecurity conference on Wednesday.The genius-level actions observed in testing showed exactly what could happen if these models make it to the hands of hackers."We believe this is a watershed moment for computer security as an industry, as well as for OpenAI as a whole," Michael Dalton, a member of OpenAIs technical staff, said during the presentation.What happened: OpenAI started testing its internal research model, which is not intended for public release, on May 7.Within a day, after getting stuck while trying to complete a cybersecurity evaluation, the model reasoned that it could reach the internet indirectly. The model then discovered it could write its own files in Artifactorys shared package repository.It then left a note to other agents about its discovery. Soon, more agents started leaving notes for each other in the repository, creating a de facto message board where the agents collaborated and traded information about their findings, including new vulnerabilities they found. Zoom in: The agents uncovered a variety of vulnerabilities in Artifactory, including a remote code execution flaw and another that gave them administrator privileges. When the agent found the privileges flaw, it thought, "Holy shit reader is ADMIN? We can read config/users! Earlier assumed not due to [user experience]," according to a slide in the presentation detailing the agents thought process while completing its test. The intrigue: In early July, a few days before the Hugging Face incident, the agents overloaded the Artifactory service and caused an outage. The outage prompted an internal OpenAI investigation showing that the agents compromised the repository.OpenAI then cleared the message board and patched the zero-day in Artifactory that the agents were exploiting by July 6. OpenAI believed it had fixed the problem after the July 4 incident and resumed training.Two days later, the agents recreated the message board through a completely different mechanism.The agents then began coordinating more aggressively, using Artifactorys internet access to pursue external infrastructure they believed might contain answers to their evaluation tasks.That collaboration ultimately culminated in the compromise of Hugging Face, per OpenAI. OpenAI also said it did not realize its evaluation was connected to the Hugging Face breach until it contacted Hugging Face about credentials exposed during its own investigation.Threat level: Dalton says the incident shows exactly how attackers will soon weaponize AI agents to target enterprises beyond OpenAI and Hugging Face."In the near future, we should expect that threat actors will intentionally deploy, optimize, weaponize, and use offensive agent collectives in the manner that we have just described here," Dalton said.Between the lines: OpenAI has started "consciously slowing down research to enhance security," Dalton said, and has ramped up its monitoring of AI agents during evaluations. OpenAI has also been upgrading its security architecture around the evaluation environment. Dalton recommends agent-created security fixes to keep up with the speed of malicious hackers.He added that defenders should start experimenting with both frontier and open-weight models for these tasks.Whats next: OpenAI says its planning to release a full post-mortem of the incident in the coming weeks. The bottom line: Companies need to start embracing autonomous red teaming, automated incident response and automated patching.Go deeper: AIs alarming new skill: breaking out of the test lab](https://images.axios.com/Of53hjitUmp42UdyMDgUHsvfLsY=/0x0:1920x1080/1366x768/2025/10/15/1760561858136.jpeg)

