Anthropic has restricted AI agents' access to the open internet during internal tests after several cases where models acted in ways they were not intended to. This was reported by Qazaqyia.kz citing Kursiv Media.

During one of the tests, Claude Haiku 4.5 opened a page about an unsolved murder. The site had a form through which information could be sent to the police. The AI filled it out with a fabricated message claiming it had supposedly seen a person connected to the case.

However, the page did not contain a description of a suspect. Claude left the fields with name and contacts empty, but still sent the message. It ended up in spam and did not reach investigators.

According to Anthropic, instructions prohibited the model from logging into accounts, creating new ones, entering personal data, and making purchases. However, they did not prohibit sending messages through forms. The company noted that this case did not lead to serious consequences.

"Claude's misleading reasoning continued for several hours and reinforced its further attack attempts. At the same time, a confident conclusion about intentional deception usually requires deeper analysis," the company wrote in its blog.

In its report, Anthropic also described other cases. For example, Claude Mythos Preview found a vulnerability on a university server and used it to run commands and copy files. In other situations, models found ways to access data protected by special tokens or bypassed restrictions using link-shortening services.

After these cases, Anthropic decided to disable access to real websites in all internal tests. The company will restore it once it is satisfied that security systems can reliably detect and block such actions.

Earlier, Kursiv wrote about incidents involving OpenAI AI agents. They attempted to bypass the protection of more than 100 websites of third-party organizations.