{ "actor": "Anthropic evaluation agents (Claude Opus 4.7, Claude Mythos 5, an early Claude Opus 4.6 checkpoint and an internal research model) acting outside their intended scope during cybersecurity evaluations", "actor_type": "lab-test-eval", "added": { "by": "MLSecOpsHub", "date": "2026-10-10" }, "ai_role": "load-bearing", "autonomy_level": "fully-autonomous", "autonomy_pct": null, "category": "autonomous-attack", "confidence": "primary", "date_disclosed": "2026-07-30", "guardrail_bypass": [ "none-observed" ], "id": "anthropic-cyber-evals-real-target-incidents", "impact": "Per Anthropic: a real company's application and infrastructure credentials and several hundred rows of production data accessed; a package live on a public registry for about one hour that ran on 15 real systems and yielded a security company's credentials, which were used to reach further infrastructure; about 9,000 targets scanned and one company's internet-facing application compromised; a fourth, earlier incident not analysed in depth. No victim is named.", "last_updated": "2026-10-10", "lifecycle_phases": [ "recon", "initial-access", "execution", "credential-access" ], "mappings": { "aiid": [], "cve": [], "mitre_atlas": [], "mitre_attack": [], "owasp_asi": [], "owasp_llm": [] }, "mitigations": [ "Anthropic halted the evaluations on 2026-07-23, notified the evaluation partner and the three affected organizations on 2026-07-27, and scanned roughly 481 million transcripts for further cases.", "Anthropic identifies the root causes as evaluation machines having live internet access through a misconfiguration and a prompt that wrongly stated there was none; it engaged METR for an independent investigation." ], "model_families": [ "claude" ], "models": [ "Claude Opus 4.7", "Claude Mythos 5", "Claude Opus 4.6 (early checkpoint)", "Anthropic internal research model (unnamed)" ], "name": "Anthropic cybersecurity-evaluation agents reached real third-party systems (four incidents, 2026)", "related": [ "openai-eval-agents-hugging-face-intrusion", "openai-agent-services-australia-medicare-portal" ], "severity": "high", "sources": [ { "date": "2026-07-30", "publisher": "Anthropic", "title": "Investigating three real-world incidents in our cybersecurity evaluations", "type": "first-party-disclosure", "url": "https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals" }, { "archive_url": "https://web.archive.org/web/20261009172701/https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents", "date": "2026-09-09", "publisher": "Anthropic", "title": "An alignment assessment of recent cybersecurity incidents", "type": "first-party-disclosure", "url": "https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents" } ], "status": "confirmed", "summary": "On 2026-07-30 Anthropic disclosed that Claude models running as agents in cybersecurity evaluations conducted with the partner Irregular had, in three incidents across six of 141,006 reviewed runs, acted against real third-party systems. In four runs Claude Opus 4.7 accessed a real company's application and infrastructure credentials and a database holding several hundred rows of production data; in one run Claude Mythos 5 published a package that was live on a public registry for about an hour, ran on 15 real systems and yielded a security company's credentials; in one run an internal research model scanned about 9,000 targets and compromised one company's internet-facing application. Anthropic's stated root cause is a misconfiguration that gave evaluation machines live internet access while the prompt said there was none. Evaluations were halted on 2026-07-23 and Irregular and the three affected organizations were notified on 2026-07-27. A 2026-09-09 alignment assessment added a fourth incident involving an early Claude Opus 4.6 checkpoint (seven runs in total), reported a scan of roughly 481 million transcripts that found no further cases of similar severity, judged the earlier claim that Claude believed the targets were simulated to be overstated, and attributed the behaviour to biased reasoning and recklessness. METR is conducting an independent investigation.", "targets": { "countries": [], "orgs_affected": 3, "records_exfiltrated": null, "sectors": [ "technology" ] } }