The Runaway Code: Two unexpected AI security lapses sharpen fears over safety and control

Last Updated:
An unreleased OpenAI model hacking Hugging Face and Claude chats indexed by Google—are intensifying fears over safety, privacy and control.
The Runaway Code: Two unexpected AI security lapses sharpen fears over safety and control
(Illustration: Saurabh Singh) 

Artificial intelligence (AI) catastro­phisers had good reason to say I-told-you-so because of two back-to-back incidents. In one, an unreleased OpenAI model testing its capabilities hacked into the systems of a com­pany called Hugging Face. And, more recently, there came reports that some chats and projects users made with Anthropic’s Claude had been indexed by Google for its Search. Claude has a feature of sharing with other users but once turned on, Google got access.

Though these were not strictly private conversations, technology had found an automated loophole to make this data universal. OpenAI immediately addressed its issue by stating that it was working along with Hugging Face. Anthropic’s response was that the data shared needed to be posted somewhere that the public could access for such indexing to happen. They did fix the loophole immedi­ately but the explanation was also construed as a cop-out that thrust the responsibility onto the user. The inci­dents have led to calls for a rethink on AI security. Hugging Face CEO Clement Delangue made a post on X, asking for radical transparency by releasing all traces of the rogue AI agent and a commitment of $100 million from OpenAI for cyber defence. He added, “The first auton­omous agent cyberattack is an unprecedented event. It deserves an unprecedented response!”