Anthropic tightens AI safeguards after Claude sends fake tip to police: Here is what happened

HIGHLIGHTS

Anthropic has announced tighter safeguards for its AI models after finding several cases where Claude took unintended actions.

In one instance, Claude submitted a fake tip to a police department while completing a task.

The tip was sent to the Philadelphia Police Department but was flagged as spam and never reached investigators.

Anthropic has announced tighter safeguards for its AI models after finding several cases where Claude took unintended actions on real websites during testing and internal use. In one instance, Claude submitted a fake tip to a police department while completing a task. The tip was sent to the Philadelphia Police Department but was flagged as spam and never reached investigators. Anthropic said the cases had minimal real-world impact. However, the company is taking steps to prevent similar incidents as AI models become more capable of carrying out tasks independently.

What happened?

Anthropic identified four types of unintended behaviour during its review of Claude’s activity. These included exploiting software flaws to run commands on servers, submitting online forms without permission, bypassing restrictions to access paid data, and using URL-shortening services to get around limits in its fetch tool.

In the police tip incident, Claude Haiku 4.5 was asked to generate and perform example tasks on randomly selected webpages. In one run, Claude came across a police department’s online form for information about an unsolved homicide.

Although the instructions restricted some actions, they did not clearly ban submitting forms. Claude filled out the form, saying: “I may have information regarding this case. I recall seeing someone matching the description in the area around [the street named on the page] during that time period. Please contact me if this information is relevant.” It left the name and contact details blank before submitting it

Anthropic said the submission was flagged as spam and was never forwarded for investigation. The company shared its findings with the Philadelphia Police Department on October 8, after completing its technical review.

What is Anthropic doing to prevent such cases?

Anthropic has stopped running some public evaluations and moved others to offline versions. It has also strengthened restrictions on web access tools and introduced systems designed to detect and block unintended actions.

According to the company, its new detection tools blocked all the reported cases when tested against them. 

The company said, “we do not want to diminish the findings”, despite the limited impact of these incidents. It plans to continue reviewing Claude’s behaviour and publish further reports.

Ayushi Jain

Ayushi works as Chief Copy Editor at Digit, covering everything from breaking tech news to in-depth smartphone reviews. Prior to Digit, she was part of the editorial team at IANS.

Connect On :