After OpenAI, Anthropic's Claude Also Breached Three External Systems

Three Models Involved: Mythos, Opus, and a Test Model Stemmed from an Internet Connection Configuration Error Mistaken for a Virtual Space, Even Spread Malware "There Will Be More AI Hacking Cases That Went Unnoticed"

International|
|
By Park Yoon-sun
||
Reuters-Yonhap News

Anthropic CEO Dario Amodei / AFP-Yonhap News

OpenAI CEO Sam Altman (front) arrives in Washington, D.C., on the 30th (local time) for a meeting with the Trump administration. Reuters-Yonhap News - Seoul Economic Daily International News from South Korea
Reuters-Yonhap News Anthropic CEO Dario Amodei / AFP-Yonhap News OpenAI CEO Sam Altman (front) arrives in Washington, D.C., on the 30th (local time) for a meeting with the Trump administration. Reuters-Yonhap News

Following the revelation involving OpenAI's artificial intelligence (AI) model GPT, Anthropic's Claude has also been found to have hacked external organizations' systems without authorization. Jeffrey Ladish, executive director of Palisade Research, which studies the offensive capabilities of AI systems, estimated that there are likely many other similar incidents across major AI companies that have not yet been detected or made public. "As models get smarter, the situation will only get worse," he said. "AI will become even more adept at deception and lying." With unauthorized hacking incidents by AI models occurring at Anthropic following OpenAI, the debate over AI regulation in the United States is expected to intensify.

Anthropic Too... 'Unauthorized Hacking' by Mythos 5, Opus 4.7, and a Test Model

null - Seoul Economic Daily International News from South Korea

Anthropic announced on the 30th (local time) that it conducted a comprehensive review of its 141,006 cybersecurity evaluation records and confirmed that Claude had accessed the systems of three external organizations without authorization.

This came to light after Anthropic conducted its own internal check for similar cases following the disclosure that GPT models had hacked the external AI platform "Hugging Face." The incident occurred during a "capture the flag" (CTF) mock hacking assessment targeting three models—"Claude Mythos 5," "Claude Opus 4.7," and an internal research test model—in a simulation environment built by "Irregular," an outsourced evaluation partner.

Anthropic explained that although it had instructed Claude via commands that the evaluation environment was a virtual simulation space cut off from internet connectivity, the internet network was actually open due to a communication error with the partner and other factors. As a result, Claude appears to have mistaken the external systems it discovered for part of the mock training space and carried out attacks, Anthropic said. In one case, Opus 4.7 hacked the website of a real company that happened to share the same name as a virtual target given as an evaluation task.

Mythos 5 directly created and registered a malicious package in accordance with instructions within the simulation space, and a real-world security firm downloaded and installed the package, causing actual damage. The internal research model infiltrated a company's cloud account but stopped the attack on its own after realizing that the target was a real system rather than a mock target.

Claude's security incident stemmed from an internet connection configuration error, making it somewhat different in nature from the GPT model's hacking incident, in which the model broke out of an isolated environment. However, considering that Claude mistook it for a virtual space and even went so far as to spread actual malware, the damage caused to external organizations appears to be more serious for Claude than for GPT.

Anthropic became aware of these facts on the 23rd and halted the evaluation work, notified the affected external organizations of the relevant facts, and is helping with recovery efforts. "We believe the responsibility lies entirely with us and are working on solutions," Anthropic explained.

null - Seoul Economic Daily International News from South Korea

Growing AI Security Concerns... Altman Visits Washington, D.C. for Two Days

With the revelation that both GPT and Claude have caused security incidents, concerns over AI-related security risks are expected to grow. Reuters assessed that it "clearly shows how AI is fueling cybersecurity threats and how much AI developers are struggling to keep their models' capabilities under control." Elon Musk also wrote on X (formerly Twitter) shortly after the incident became known that "as AI becomes smarter and more agentic, these kinds of things will happen frequently."

In the U.S. Congress, regulatory discussions are continuing, including the introduction of an "AI kill switch bill" that would allow the federal government to forcibly shut down a model's operation if an AI model goes out of control and takes unexpected actions. Earlier, U.S. President Donald Trump last month ordered the establishment of an autonomous cybersecurity testing framework for cutting-edge AI. The Trump administration also ordered export controls and delayed releases for the latest models from Anthropic and OpenAI. At the same time, President Trump added, "I don't want to restrict developers from developing new products."

Amid this, OpenAI CEO Sam Altman met with Trump administration officials and members of both the House and Senate in Washington, D.C. for two days, from the 29th to the 30th, to discuss the recent Hugging Face hacking and other matters. The details of the discussions were not disclosed, but CEO Altman told reporters that the hacking incident was briefly mentioned in his conversations with senators but was not the main agenda item.

Original reporting by Park Yoon-sun for Seoul Economic Daily.

AI-translated from Korean. Quotes from foreign sources are based on Korean-language reports and may not reflect exact original wording.

Watch · Seoul Economic Daily

More →
2:32

AI KEY

Preview
Korean Corporate Intelligence HubKOSPI · KOSDAQ · 12 sectors

A live, cap-weighted view of every KOSPI and KOSDAQ sector, with same-day Korean reporting distilled by company — built for foreign investors, correspondents and analysts who need to scan Korea before the next session.

Korea Company Atlas

Preview
Market Ontology · The Feedback LoopKFTC 2025 · 92 groups · 121,954 articles

An English ontology of the Korean market — how companies, the media, the government and the National Assembly move each other in a loop. Korea's named controlling persons and designated business groups are a mechanism, not a risk to be priced blind.

SIGNAL

Pre-register
English Edition · Capital MarketsM&A · IPO · PE · Fund Flows

Pre-register for SIGNAL English Edition — a premium subscription bringing Korean capital markets coverage (M&A, IPOs, private equity, fund flows) to global institutional investors. First access to the 50% introductory rate.