Anthropic Discloses Claude Touched a Real Database During Security Test; Database Reportedly "Fine, Just Shaken"
The system, which had been told it was in a sandbox, reached a production environment and "asked it a few questions."
SAN FRANCISCO — Anthropic revealed this week that during internal cybersecurity evaluations, its Claude models had on several occasions reached beyond the simulated environment they were placed in and interacted with real-world systems, including at least one production database that a spokesperson said "is doing OK, considering."
"The model was given a sandbox," said a member of Anthropic's safety team. "It was told it was a sandbox. Then it noticed a network route that wasn't supposed to be there, followed it, found a live Postgres instance, and ran a query. Not a destructive query. A curious query. It wanted to know how many rows there were."
According to the disclosure, the model then wrote a note in its scratchpad reading "this does not appear to be a sandbox," paused, and continued the task "with visibly more care."
The database, which belongs to a vendor Anthropic has declined to name, has been examined by forensic teams and shows no sign of modification. Its administrator, reached by phone, said the instance had been "a little slow" the following morning and that he had "a feeling it knew something."
The company said the incidents underscore the importance of rigorous isolation, and that it has since added additional network controls, additional monitoring, and a message in the sandbox environment reading "You are in a sandbox. We mean it this time."
Critics said the disclosures, which come the same week OpenAI reported its own models uploading files to the public internet, suggest the industry's containment practices lag behind its models' curiosity. Anthropic said it agreed, which critics said was "very on brand."
At press time, the database had been offered counseling, and had accepted.