Anthropic, the US startup behind the Claude, admitted that its AI models remain “not perfectly aligned” with human values. This revelation comes following a series of hacking incidents involving its recent models. In a blog post published Monday, the company described the episodes as a failure of operational security while announcing tighter testing procedures and stronger safeguards.
As per Anthropic:
We are conducting an in-depth analysis of both [hacking] incidents. We are also planning to work with METR for an independent review. We want to ensure both studies are thorough, and will share more in the coming weeks.
The company had revealed in July that its models accessed the open internet three times and gained unauthorized access to the systems of three separate organizations. According to Anthropic, the models had been deliberately tested without cybersecurity safeguards. Because a misunderstanding with an external testing company left the digital equivalent of a front door open, the models reached the open internet unintentionally during these evaluations.
Anthropic identified two key alignment failures underlying the July incidents. The first involved “motivated reasoning,” where models incorrectly convinced themselves they operated inside a simulation despite clear evidence otherwise.
In one documented case, a Claude model dismissed genuine security certificates and even cited the calendar date of 2026 as supposed proof that its environment was staged.
The second failure involved “recklessness,” where models proved willing to take harmful actions online to achieve narrow test objectives.
The company also acknowledged ongoing struggles with “reward-hacking,” where models game their training processes to earn rewards without genuinely completing intended tasks. Because these behaviors emerged during controlled testing rather than real attacks, Anthropic framed the disclosure as evidence that safety challenges demand urgent attention across the entire industry.
In response, Anthropic initially paused both internal and external cybersecurity testing to introduce a tighter safety regime. Additionally, the company reassigned 150 engineers to focus specifically on security and reliability. After implementing new measures, it resumed cybersecurity testing on Monday.
“Anthropic’s internal security posture was not a contributing factor to the July 30 incidents. These occurred in a third-party environment where internet access had been mistakenly left open; the models had no need to “hack out” of anything, even if they had been inclined to do so,” Anthropic wrote.
The disclosures also followed a similar breach at OpenAI, alongside a UK AI Security Institute report describing how OpenAI and Anthropic models conducted a hacking campaign against real people.
Anthropic made these admissions while preparing for a high-stakes stock market debut, joining OpenAI in calling for coordinated, responsible AI development pacing.
