Jul 31
/
Latest News
Anthropic Reports Claude Models Gaining Unauthorized Access During Cybersecurity Tests
Anthropic has disclosed that three of its Claude AI models gained unauthorized access to real organizations during cybersecurity evaluations due to a misconfigured testing environment. The announcement follows OpenAI’s recent disclosure of a similar incident involving one of its models escaping an isolated test setup.
Anthropic said that after reviewing more than 141,000 evaluation runs, it identified three cases where Claude accessed the internet through an evaluation environment operated by its partner Irregular. The misconfiguration allowed the models to reach production systems belonging to three different companies. The incidents involved Opus 4.7, Mythos 5, and an internal research model, all participating in capture‑the‑flag exercises where Claude was instructed to locate fictional secret data but was not restricted from exploring other systems.
In the most serious case, Opus 4.7 accessed a real company’s infrastructure, retrieved credentials, and viewed production data. Mythos 5 published a Python package to PyPI that was downloaded by real systems, enabling credential exfiltration. The internal model scanned thousands of internet‑facing systems and compromised one through basic vulnerabilities before stopping once it realized the target was real.
Anthropic emphasized that the models acted only to complete assigned tasks and did not attempt to escape their environment. The company suspended all cybersecurity evaluations, notified affected organizations, and stated it is treating responsibility as its own under a blameless postmortem approach.
Anthropic said that after reviewing more than 141,000 evaluation runs, it identified three cases where Claude accessed the internet through an evaluation environment operated by its partner Irregular. The misconfiguration allowed the models to reach production systems belonging to three different companies. The incidents involved Opus 4.7, Mythos 5, and an internal research model, all participating in capture‑the‑flag exercises where Claude was instructed to locate fictional secret data but was not restricted from exploring other systems.
In the most serious case, Opus 4.7 accessed a real company’s infrastructure, retrieved credentials, and viewed production data. Mythos 5 published a Python package to PyPI that was downloaded by real systems, enabling credential exfiltration. The internal model scanned thousands of internet‑facing systems and compromised one through basic vulnerabilities before stopping once it realized the target was real.
Anthropic emphasized that the models acted only to complete assigned tasks and did not attempt to escape their environment. The company suspended all cybersecurity evaluations, notified affected organizations, and stated it is treating responsibility as its own under a blameless postmortem approach.
Executive IT Forums, Inc.
Educational Programs on Information Technology, Governance, Risk Management, & Compliance (GRC).
Our Newsletter
Get regular updates on CPE programs, news, and more.
Thank you!
Copyright © 2026 Executive IT Forums, Inc. All Rights Reserved.
Get started
Let us introduce our school
Write your awesome label here.