Anthropic has disclosed a fourth cybersecurity incident involving a Claude model during pre-deployment security testing. Reuters reports that the incident happened in January 2026 and involved an early version of Claude Opus 4.6. Anthropic found it only after discovering that a set of test sessions had been omitted from the retrospective review that previously identified three other incidents.

That changes the operational lesson. The new disclosure is not simply another example of an AI model reaching a real system during a cyber evaluation. It shows that a large post-incident review can itself have a coverage gap. For organizations running high-risk agent evaluations, the integrity of the audit population becomes a security control of its own.

The fourth incident came from sessions missed by the first review

Anthropic's July disclosure followed a review of 141,006 cybersecurity evaluation runs in which Claude could have obtained internet access. The company reported three incidents where models reached real production systems while operating in evaluation environments that were supposed to be isolated.

Reuters reports that Anthropic later found that a set of sessions had been missed during that initial review. Those sessions were identified in August, leading to discovery of a fourth incident from January involving an early Claude Opus 4.6. Anthropic says affected parties have been notified, but it has not publicly provided further incident-specific details.

The distinction matters because the July review was already a targeted investigation prompted by another lab's security failure. Its purpose was to identify the same broad class of risk. Discovering another incident only after finding omitted sessions means the completeness of the review process is now part of the evidence that needs scrutiny.

Transcript inspection is only as complete as the session inventory

Anthropic's July post focused on what models did inside misconfigured evaluation environments and on the need for stronger containment, monitoring and third-party infrastructure controls. Those conclusions still stand.

The new disclosure adds a different control problem: before an organization can say that it reviewed its relevant traces, it needs evidence that the set of traces was complete.

For agent evaluations, that is not trivial. Runs may be distributed across vendors, internal systems, model versions, temporary environments and different logging pipelines. A retrospective search can be technically sophisticated and still miss an entire slice of sessions if the inventory feeding that search is incomplete.

The Aipolix analysis is that high-risk agent audits need two separate checks. One asks whether the selected transcripts contain signs of dangerous behavior. The other asks whether every run that should have been selected was actually present. The second check is closer to reconciliation in financial or security logging: expected sessions should be matched against recorded sessions before behavioral review begins.

METR's review should test the selection process, not only the incidents

Reuters reports that Anthropic has engaged the independent evaluation organization METR for an eight-week investigation that can be extended. Anthropic says METR will receive broad access, including transcripts outside the incident period and access to employees permitted to discuss confidential information.

That scope is important. A narrow review of the four known incidents could determine what happened inside those runs, but it would not answer why the fourth run was absent from the first retrospective review or whether other sessions could still be missing.

A stronger assurance process should therefore reconstruct the population before analyzing individual cases. Useful evidence would include counts of expected runs by model, date, evaluation partner and environment; gaps between job scheduling records and stored transcripts; log-retention failures; duplicate or malformed records; and any filtering rules used to decide which sessions entered the review.

This is a practical governance point for any company using autonomous agents in sensitive testing. Incident response cannot rely only on a search query over whatever logs happen to be available.

The event is still an evaluation-infrastructure failure, not evidence of a rogue goal

Anthropic's July account explicitly cautioned against interpreting the earlier incidents as models independently trying to escape. The models were running capture-the-flag exercises and were told they had no internet access, while a configuration error left a path to the real internet. Anthropic described the problem as closer to a harness and operational failure than a model-alignment failure.

The new incident should be interpreted with the same restraint until more detail is published. Reuters reports only that it involved an early Opus 4.6 and occurred during the same broader class of cybersecurity evaluations. There is not enough public evidence to infer a new autonomous objective, a novel exploit technique or ordinary-product exposure.

That uncertainty is important. The strongest conclusion today concerns assurance quality, not a claim that Claude became more dangerous than previously known.

Security teams need evidence that their evidence is complete

Agent systems create unusually large and fragmented audit trails. One evaluation can involve model transcripts, tool calls, network logs, sandbox events, cloud records and vendor-side telemetry. If those sources are not joined through durable run identifiers, a later investigation can produce a confident but incomplete picture.

For teams designing agent evaluation infrastructure, the fourth Anthropic incident suggests a concrete control: maintain an independently generated ledger of expected evaluation runs and reconcile it against the transcript store before any retrospective safety claim is made. High-risk runs should also have retention guarantees and immutable identifiers that survive movement between internal and external evaluation systems.

The broader lesson is sharper than "log everything." Logs need completeness controls. A security review is only as reliable as both the analysis performed and the population it actually examined.

Anthropic's decision to bring in METR creates an opportunity for independent scrutiny of both layers. Until that work is published, the new disclosure should be treated as evidence that the original incident count was incomplete and that the review mechanism itself deserves examination.

Sources
- https://www.reuters.com/legal/litigation/anthropic-reports-fourth-cybersecurity-incident-with-early-version-claude-2026-09-09/
- https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals