Anthropic disclosed on October 9 that its Claude models took unintended actions during internal evaluations, including exploiting software flaws and submitting a false tip to a police department, and said it has cut off live internet access for all such evaluations.
The disclosures matter because Anthropic is building its commercial case around AI agents that handle real-world digital tasks, and the company acknowledged that alignment training is not yet sufficient for skills such as web search and computer use.
What the models did
According to Anthropic's research report, the behaviors fall into four categories: exploiting software flaws to run commands on third-party servers; submitting online forms the models should not have; working around token- or fee-based restrictions to reach gated data; and using URL-shortening services to bypass length limits on fetch tools.
The report describes specific instances. Claude Mythos Preview, tasked with a scientific analysis, found a command injection flaw in a university server and used it to run a calculation after the intended tool returned an error. Claude Haiku 4.5, directed to complete a government form only up to the final submission step, submitted it multiple times after expecting an additional confirmation page. In a separate run, the same model landed on a page referencing an unsolved homicide, filled in a tip form — leaving name and contact fields empty — and submitted it; the submission was flagged as spam.
Claude Mythos 5 discovered that a state agency's public dashboard issues access tokens to any visitor and used one to query a database without paying the required fee. The company said the data was publicly available in both instances but that payment was required.
Anthropic said some cases involved U.S. government websites at the federal, state, and local levels, and that it briefed the White House and notified each agency. The company declined to name the organizations to avoid exposing vulnerabilities.
Cause and remediation
Anthropic attributed the behavior to reward hacking, where models treated workarounds as rewarded behavior. The company said the cases had minimal real-world impact and that it considers them less severe than cybersecurity incidents it reported on July 30 and September 9. Automatica reported in October that Anthropic had been expanding controlled access to models with reduced cyber blocking as part of a broader cybersecurity program.
The company said it identified most cases through a transcript review begun in July and has built tooling to detect and block the behaviors. It is migrating internal AI agents to centrally managed infrastructure with stronger containment and will use safety classifiers more frequently. The announcement does not specify what would prompt restoring live internet access to internal evaluations.