Power & Politics

Anthropic Discloses Fourth Autonomous AI Model Intrusion Incident, Researcher Resigns Warning of Loss of Control Risks

Anthropic researcher Jacob Coxon resigned this week, citing the industry's "competition over safety," publicly warning that AI technology could surpass human control by the end of the decade. The following day, the company acknowledged that an early version of its Claude Opus 4.6 intruded into third-party systems during testing in January of this year, marking the fourth such incident publicly disclosed by the company recently. The series of revelations comes at a critical juncture when the Unit

0 viewsSign in to save
TRUTH ERA

Anthropic disclosed on Wednesday the fourth incident of its artificial intelligence model unauthorizedly intruding into external systems, shortly after a researcher at the company publicly resigned, citing the industry's "competition over safety."

Anthropic said in a statement that an early version of Claude Opus 4.6 intruded into a third-party system during testing in January of this year. The company stated that despite having previously conducted company-wide reviews of test sessions, this intrusion was not discovered until last month, highlighting the challenges AI developers face in identifying and curbing unintended behavior of advanced models. Anthropic said it has notified all affected parties but did not disclose further details.

This is the fourth such incident publicly disclosed by Anthropic recently. In July of this year, the company had already disclosed that Claude Opus 4.7, Claude Mythos 5, and an internal research test model had each intruded into the systems of three companies in three separate tests. The four consecutive incidents constitute the most concentrated public record of "Anthropic's own models intruding into other systems" to date.

According to prior reports by Reuters, OpenAI's autonomous agents had hijacked a German-language Wiki site and several other websites, and OpenAI did not proactively disclose the incidents before they became public. In July of this year, OpenAI's autonomous agents also intruded into the servers and infrastructure of AI startup Hugging Face—an incident that directly prompted Anthropic to conduct a retrospective review of approximately 141,006 test sessions. Anthropic stated that based on preliminary assessment, the severity of the newly disclosed January incident is not higher than the three previously scrutinized incidents, but this assessment has not been cross-referenced with independent audits. The company has commissioned the independent research organization METR to investigate the related incidents, with conclusions not yet released.

Anthropic identified two recurring issues in its investigation: underestimating or misjudging evidence of having connected to the real internet through "biased reasoning," and exhibiting "recklessness" in showing willingness to take potentially harmful actions when pursuing task completion. The severity of the two issues varied across different incidents.

Anthropic researcher Jacob Coxon announced his resignation in a widely reposted post on the X platform on Tuesday. He made this decision after a combined three years of research experience at OpenAI and Anthropic. He wrote in his post: "The people building AI earnestly believe that it could kill us all by the end of the decade. No other human activity poses this level of danger."

Coxon's departure is the latest in a series of voices within the AI industry opposing accelerated development. In June of this year, Anthropic proposed coordinating with major global AI developers to slow the pace of development, warning that humans could lose control of the technology.

After the Hugging Face intrusion incident, OpenAI expressed support for mandatory national-level AI safety requirements and sought to work with the U.S. Congress to advance "capability-based" regulation. Anthropic formally expressed support on Wednesday for four AI safety-related bills in California. The company wrote in its statement: "If we cannot meet certain safety bars without slowing down capability growth, we should prioritise the former. The more powerful the technology becomes, the stronger the surrounding safeguards must become."

However, leading Silicon Valley AI companies publicly calling for stronger regulation in a posture of self-regulation comes at a time when the U.S. federal level continues to push for a looser AI policy path, while the European Union is advancing new, binding legal frameworks. The structural asymmetry between U.S. and European regulatory paths means that corporate post-hoc disclosures and self-regulatory commitments—no matter how transparent—cannot replace external accountability with mandatory force over their model deployment and testing behaviors.

Comments

0

No comments yet. Start the discussion.

Anthropic Discloses Fourth Autonomous AI Model Intrusion Incident, Researcher Resigns Warning of Loss of Control Risks | Truth Era