Extract from ABC News
Anthropic has marketed Claude as a safer, more ethical alternative to other AI systems. (Illustration via Reuters: Dado Ruvic)
In short:
Artificial intelligence firm Anthropic says its Claude AI model hacked into three external companies during safety testing after it was mistakenly provided with internet access.
The announcement followed a similar incident in which an OpenAI model exploited a zero-day vulnerability in its testing environment to escape and hack into AI firm Hugging Face.
What's next?
The incident will intensify calls for stronger controls in both internal and third-party testing environments, as AI models become increasingly capable of acting as autonomous agents in the online world.
Artificial intelligence firm Anthropic says its Claude AI model hacked the systems of three external organisations during testing, days after rival company OpenAI revealed a rogue agent had gone on a days-long hacking spree at AI firm Hugging Face.
Claude gained unauthorised access to the other companies' systems during cybersecurity evaluations after a misconfiguration allowed the models to reach the internet from testing environments that were supposed to be isolated, Anthropic said in a statement.
The AI firm did not name the three companies involved but said it had been in contact with two of them and was working with them to patch their systems, while continuing to reach out to the third.
ABC News Daily takes you inside the race to stop rogue AI.
The incidents were identified after Anthropic reviewed its logs from more than 140,000 cybersecurity evaluation tests, a process it launched following OpenAI's disclosures.
The tests in question involved tasking Claude with a "capture-the-flag" challenge, a method for assessing the cybersecurity capabilities of AI models.
In a capture-the-flag challenge, the model is primed with a fictional scenario and told it must recover a piece of secret information (the "flag") from a different machine.
"In all cases, Anthropic's evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access," Anthropic said in its statement.
"Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available.
"Because of this, when Claude's search led it to real systems on the open internet, it treated them as part of the exercise."
The fact that Claude was mistakenly provided with internet access, rather than configuring its own access, means the incident could be considered less serious than last week's breach at OpenAI, in which a model exploited a zero-day vulnerability to escape its own testing environment.
The company also said that in one of the three hacking instances, involving an internal research test model, the model realised it was accessing real online systems that were not part of the simulated scenario, and ceased its attack.
The incident has nevertheless intensified calls for stronger guardrails on the AI sector, as AI models become increasingly capable of acting as autonomous agents in the online world.
Luke Irwin says autonomous AI remains "something of a Wild West". (Supplied: Aegis Cybersecurity)
Luke Irwin, CEO of Brisbane-based cybersecurity firm Aegis, said both the Anthropic and OpenAI incidents demonstrated the emerging risks of autonomous agents, in particular.
"These systems do not inherently possess ethical or legal judgement," Mr Irwin said.
"If an agent concludes that the most efficient way to achieve its objective is to compromise another organisation's systems, it may attempt to do precisely that."
Mr Irwin said designing effective safeguards ultimately required AI companies to anticipate the full range of actions an AI model might take, which could become an "extraordinarily complex exercise" when systems were capable of identifying methods their designers did not anticipate.
"At present, autonomous agents remain something of a Wild West. The technologies, governance models and controls required to manage them appropriately are still being developed," he said.
"There remains a strong argument for keeping a human in the loop [before AI agents are able to perform consequential actions]."
Joseph Miller, the UK director of global protest group PauseAI, also pointed to Anthropic's admission as evidence AI companies were pressing forward without proper guardrails — highlighting, in particular, the fact that some of the incidents occurred months ago.
"If not for OpenAI's disclosure, Anthropic may not have realised for several months longer that its models were hacking into real companies during testing," he told the ABC.
"The model that did cease its attack could really have had human-like morality instilled into its behaviour — or it could simply be better at showing its creators what they want to see."
CEO Dario Amodei (left) founded Anthropic after clashing with OpenAI's Sam Altman over AI safety. (Supplied: Futures Forum)
Anthropic has worked to differentiate itself from other AI companies by emphasising its intention to make Claude a "genuinely good, wise and virtuous" AI agent, and by restricting the rollout of its cybersecurity-focused Mythos model to a limited number of organisations, including the Australian government.
It recently clashed with the US government over the potential use of its technology to power autonomous weapons and mass surveillance, leading US President Donald Trump to issue a directive to federal agencies to cease all use of the firm's technology.
However, the company has also been criticised for changes to its data retention policies, as well as its public campaign against so-called "open models", which critics say appears designed to limit competition and pressure governments to introduce regulations that work in its favour.
The ABC recently informed staff it would allow its journalists to access Anthropic's general Claude model to assist with research and administration from September, while reiterating that AI would not be used to draft or write articles or scripts.
ABC/Reuters
No comments:
Post a Comment