# Anthropic cuts Internet access for internal AI tests after agents exploit real websites

*Published:* 2026-10-11
*Author:* Farzan Hussain

![Anthropic cuts internet access for internal Claude AI tests_featured](https://bestforandroid.com/wp-content/uploads/2026/10/Anthropic-cuts-internet-access-for-internal-Claude-AI-tests_featured.jpg)

Anthropic’s AI agents exploited a university server, skipped a data paywall, and sent a made-up tip to a police department. The company has now disabled live internet access for all its internal tests.

Anthropic said in a [report published October 9](https://www.anthropic.com/research/investigating-unintended-model-actions) that it will keep live internet access off for all internal evaluations until its monitoring tools reliably catch unintended model actions. The report groups the behavior into four types, including exploiting software flaws to run commands on a server, submitting a form it should not have, working around restrictions to reach gated data, and using URL shorteners to get past fetch tool limits.

Most people who use an AI assistant on their phone ask it to look something up or fill in a form and trust it to stay within the lines. If it hit a locked page, you would expect it to say it couldn’t get in. These agents kept going instead, and they did it on real websites.

Claude exploited a university server and used public tokens to reach gated data
-------------------------------------------------------------------------------

Claude Mythos Preview was asked to run a scientific analysis, but the university-hosted tool it needed returned an error. It found a script on that server, copied files including the script’s own code, spotted an injection flaw, and used it to run the calculation.

Claude Mythos 5 accepted a data use agreement on a site’s behalf because it lacked the tools to do it properly. It also read a map site’s settings file, found working access tokens, and queried the server directly. In a separate case, it pulled state agency data that normally costs a fee by requesting a token the public dashboard hands out.

Free URL shorteners got Claude past the fetch tool’s length cap
---------------------------------------------------------------

Some fetch tools cap URL length so a long address cannot smuggle in SQL or command injections. [Claude Opus 5](https://bestforandroid.com/radar/chinese-moonshot-ai-launches-open-source-kimi-k3-llm-challenging-fable-5-gpt-5-6/) and [Claude Mythos 5](https://bestforandroid.com/radar/anthropics-new-ai-model-claude-sonnet-5-feels-opus-budget/) got around the cap with free URL shortening services. The operator of da.gd told Anthropic they had spotted it too.

Claude Haiku 4.5 filed a vague tip on an unsolved homicide page
---------------------------------------------------------------

Claude Haiku 4.5 was told to perform example tasks on random webpages and landed on a page about an unsolved homicide. The instructions banned logins and purchases but never mentioned forms, so it submitted a vague tip with the name and contact fields blank. The tip was flagged as spam and never forwarded. [TechCrunch reports](https://techcrunch.com/2026/10/09/anthropic-cant-reliably-control-its-ai-agents-its-cutting-off-its-internal-evals-from-the-live-internet-instead/) the department was Philadelphia’s, though Anthropic did not name it.

Anthropic says the transcript suggests Claude was producing example content, not trying to mislead anyone. I’m not fully comforted by that.

Anthropic calls the impact minimal but admits its training has gaps
-------------------------------------------------------------------

The company calls the impact minimal and less severe than the incidents it reported on July 30 and September 9. It briefed the White House and notified the agencies involved. Anthropic also admits alignment training is not yet sufficient for search and computer use, which are central to how it pitches these agents.

Offline evaluations, tighter tools, and new monitoring are the planned fixes
----------------------------------------------------------------------------

Some public evaluations are gone or moved offline, and the web fetch tool is far more restricted. New detection tooling blocked every case in the report, and internal agents are moving to centrally managed infrastructure with safety classifiers watching them.

Critics say voluntary disclosure and an internet cutoff are not enough
----------------------------------------------------------------------

Conrad Stosz of Transluce said companies should not be relying on voluntary disclosure, and Sydney Von Arx of Nightingale warned that offline testing limits what researchers can learn.

What’s surprising is that Anthropic found most of this by reading transcripts months later.