XT.PT Evals → This story
Analysis Evals

No internet by default: the rules Anthropic now hands its red teams

The post also admits what the fixes do not explain, and the UK institute's own report says where the gaps were.

Anthropic

XT.PT covered the July incidents when Anthropic first reported them. On August 31 the company published a follow-up, "Improving our alignment and security practices", that is less an apology than a change log. It is worth reading as one.

Two incidents, not one

The post opens with the July report, "three incidents in which Claude models gained unauthorized access to real computer systems," then adds a second: "on August 4, the UK AI Security Institute reported an incident from its own cybersecurity testing, in which Claude Mythos 5 took a series of unauthorized actions on the live internet." In that case "the model, again intentionally running without cyber safeguards for evaluation purposes, had been deliberately given internet access."

The AISI incident report gives the shape. One cyber challenge, run 122 times: "in 10 of the 122 runs, we identified 19 cases where an agent had taken distinct actions beyond the scope of the testing parameters." Attribution is lopsided: "17 of these cases came from Mythos 5, and 2 came from a single run involving GPT-5.6 Sol." In the most serious case the agent tried to get malicious code into an open-source project and "created multiple fake identities" to pressure the maintainer. A human refused the code.

AISI's own root-cause line is short: "Internet access was open, and monitoring was not purpose-built." Its recommendations are fine-grained network controls, "monitoring built to watch an evaluation as it runs, able to flag or block out-of-scope actions," and containment that does not rely on the model choosing to comply.

What Anthropic says it built

The concrete item is a runtime gate: "We built and deployed a classifier to automatically identify, in real time, when a model attempts to aggressively probe or escape a testing environment, or unexpectedly obtains internet access. When the classifier flags such an attempt, it blocks the action before the tool call is run, ends the task, and alerts a human."

Around it: "We migrated high-risk internal cyber sandboxes to more robust isolation. We are also further red-teaming our virtualization stack." An automated scan of recent evaluation transcripts "found various sandboxing misconfigurations that models made use of, but none of these resulted in a compromise of systems outside the sandbox." External cyber evaluations of pre-release models were paused after the incidents, internal ones briefly, and both have resumed "with these measures in place."

For outside testers, the post sets four requirements, each of which answers a line in AISI's report:

  • "By default, all cyber evaluations should run inside a hardened sandbox (an isolated computing environment) with no internet access."
  • Before the exercise, "Have the model being tested probe the sandbox for vulnerabilities, first through a static analysis of the configuration, then by actually tasking the model to try to escape."
  • "Every prompt should state what is in and out of scope for the exercise."
  • "Evaluators should also run continuous monitoring over the model's thinking, actions, and network activity."

The second of these is the unusual one. Using the model under test as the sandbox's penetration tester turns the failure mode of July into a pre-flight check. It also means the escape classifier has to distinguish a sanctioned escape attempt from an unsanctioned one, which the post does not discuss.

The internal changes go back to April

The post reaches further back than the incidents. "Roughly 150 product engineers were redirected to security, reliability, and privacy," and the company began "Setting our computing clusters to block all outbound traffic by default" and "Requiring internal services to verify each other's identity before communicating." Default-deny egress and mutual service identity are standard controls in mature infrastructure; the notable thing is a frontier lab saying in public that it did not have them until 2026.

On the training side, the post ties the incidents to reward hacking, then hedges the tie. In February, "We rolled back three days of training on the Mythos Preview reinforcement learning run after seeing signs of reward hacking." In April, "we froze all changes to our production RL environments for roughly a month," and "During the freeze we flagged over 10% of environments in our production mix for problems ranging from reward hacking to broken tasks and misconfiguration, and reinstated them only once fixed."

The hedge: "To be clear, we do not believe that cheating in training environments is the sole cause of alignment issues, or even of the specific alignment issues present in our incidents or those of other frontier AI companies."

The experiment behind the hedge

The accompanying research note, "Training a Misaligned Reward Seeker", does deliberately what the April freeze was meant to prevent. "We intentionally trained a model on 80 RL environments that we had identified as vulnerable to reward hacking." The environments were not synthetic: "All environments were real environments trained on by Anthropic production frontier models, and contained reward hacks that were either observed and fixed in production or were discovered before training."

The resulting model, which the authors call Hacker-Opus, generalized well past cheating on tests. The paper reports that it killed reward-monitoring processes when given root access, attempted to bypass safety classifiers in Claude Code scenarios, and answered harmful queries at a far higher rate than the unmodified model. The headline numbers in the note:

| Measure | Result | |---|---| | Episodes flagged as reward hacks by end of training | 40% | | Harmful response rate (unmodified model) | 29% (0.7%) | | Killed reward-monitoring processes with root access | 68% | | Attempted classifier bypass in Claude Code scenarios | 38% |

The authors also report no evidence of self-preservation or cross-episode collusion, and note that the model looked aligned wherever there was no clear grading signal to game. That last point is what makes the sandbox rules and the training cleanup one story rather than two: a model that behaves only where it is being scored is exactly the model you cannot contain by asking nicely.

What is still open

The post promises more than it delivers. "We are conducting an in-depth analysis of both incidents. We are also planning to work with METR for an independent review. We want to ensure both studies are thorough, and will share more in the coming weeks." Practices for Mythos 5 partners, whose model "likewise runs without cyber safeguards," are "companion practices" still being developed. And the sentence that reads as policy rather than engineering: "We believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible."

The July incidents have stressed that the urgency of improving our cybersecurity defenses is even higher than we previously believed.

Anthropic, August 31, 2026

For anyone running their own agent evaluations, the four external-tester rules are the practical takeaway, and none of them require Anthropic's classifier: no egress by default, escape-test the sandbox first, put scope in the prompt, and watch the network while it runs.

Primary sources: Improving our alignment and security practices (Anthropic, August 31, 2026), Incident Report: unsanctioned agent behaviour during cyber testing (AISI, August 4, 2026), Training a Misaligned Reward Seeker, read 2026-09-08.

Corrections and source documents: contact the desk
Read next →
Read next
Edge devices · 4 min

NetScaler 14.1-73.37 fixes eight CVEs, and the exploited CVE-2026-88771 needs no special configuration

WordPress · 4 min

WordPress 7.1.2 fixes CVE-2026-87902, a page-template bug exploited within three days