Anthropic has published a detailed report on how threat actors misused its Claude AI. The company’s Threat Intelligence team identified and disrupted these operations. The report covers activity between December 2025 and August 2026. It spans seven distinct harm areas across the globe.
As per Anthropic:
Over the past six months, our Threat Intelligence team identified and disrupted a series of cyber operations in which threat actors used Claude. The actors included suspected state-sponsored groups, financially motivated criminals, and politically motivated individuals.
These areas include cyber operations, influence campaigns, and surveillance. They also cover scams, fraud, biological misuse, and conventional weapons. Notably, the abuse involved Claude Haiku, Sonnet, and Opus models. Anthropic said its safeguarded Fable and Mythos models saw almost no misuse.
Sophisticated Attacks Without Sophisticated Attackers
A central finding concerns how AI lowers the skill barrier. Essentially, AI collapsed the gap between elite and amateur hackers. Consequently, lone individuals now run campaigns once needing whole teams. Anthropic warns sophistication is no longer a reliable attribution signal.
Among the most alarming cases, Anthropic disrupted a Yemen-based weapons engineering cell. The company said the group operated out of northern Yemen. It was reportedly running three separate weapons development programs simultaneously. These spanned guided rockets and longer-range missile ambitions.
Critically, the actors used Claude Code in place of human software engineers. Specifically, they tasked it with developing guidance and control software. This is the software that steers and stabilizes a flying vehicle. Consequently, the case shows AI potentially substituting for scarce technical expertise.
The actors also worked in a deliberately organized manner. They ran several Claude instances at once, assigning each a role. One instance wrote code while another handled research. A third reviewed the output, mimicking a small engineering team.
Anthropic stressed that its safeguards blocked many of the requests. However, it acknowledged that not all attempts were stopped. The actors used evasion tactics to circumvent those protections. Notably, they hid their true goals and split work across sessions. This prevented any single session from revealing full intent.
Importantly, Anthropic found no evidence the actors fielded a working weapon. However, they did test-fire a guided rocket at one point. That field test appears to have failed shortly after launch. The company banned the accounts and alerted relevant partners.
The report documents this shift through several striking cases. In one, a single hacktivist targeted European political parties. They used AI across the entire attack chain effectively. Remarkably, one person built a mass privacy-attack platform alone.
AI Grows More Autonomous
The report also highlights AI’s increasingly autonomous role in attacks. Many operations went beyond simple chatbot questions and answers. Instead, multi-agent frameworks handled reconnaissance, exploitation, and data theft. Humans mostly retained control over targeting and reviewing results.
One Russian espionage actor automated much of their operation. Their AI monitored whether security products detected their malware. If detected, the AI automatically rebuilt the malware to evade defenses. This effectively inverted the cost burden back onto defenders.
The AI Supply Chain Itself Is a Target
A notable trend involves attackers targeting AI access directly. Criminals increasingly steal API keys and session tokens as loot. Stolen keys grant attackers free compute and useful cover. Importantly, these keys came from customers’ environments, not Anthropic’s systems.
One group ran a fraudulent reseller offering cheap “Claude” access. However, it secretly proxied traffic and harvested users’ credentials. Another actor’s explicit goal was accessing a pre-release Claude model. That attempt ultimately failed across more than a dozen attempts.
Influence Operations and Fraud
Beyond cyberattacks, the report details nine influence operations. These originated in Russia, Iran, Turkey, and elsewhere globally. Several campaigns were deliberately timed around national elections. Actors used Claude to build fake news sites and personas.
Ultimately, Anthropic stresses it disrupted every operation described. It strengthened its safeguards using lessons from each case. The company also shared intelligence with authorities and industry partners. As models grow more capable, it warns these risks will intensify. They concluded their report on new hope lingering on new horizons:
We’ve also added new safeguards that make it harder for unauthorized labs to distill Claude’s capabilities. Claude now summarizes its internal reasoning before responding, which makes stolen transcripts less useful for training another model. And with Fable 5.1 we introduced preserved thinking, which stops new API accounts from altering the system prompt, tools, or messages that precede Claude’s reasoning in multi-turn conversations. That reasoning is encrypted, but editing the context before it is a common technique attackers use to make Claude reveal it.
