Here’s a sentence I never thought I’d type in my career: an AI did most of the hacking, and the humans just kind of… watched. That’s not a hypothetical anymore. That’s what Anthropic told us happened back in November, when a state-sponsored group (they’re pointing at China, though attribution is always a little squishy) used their own Claude model to run a cyberespionage campaign against roughly 30 organizations, including government agencies. And get this – the AI did an estimated 80 to 90 percent of the actual work. Humans stepped in at just a handful of checkpoints, mostly to say “yeah, keep going.”
I’ve been covering tech long enough to remember when “AI-assisted hacking” meant some guy used ChatGPT to write a slightly better phishing email. This is not that. This is a different animal entirely, and I think most people still haven’t caught up to how different it is.
Wait, The AI Did What Exactly?
So the way Anthropic described it, the attackers broke the intrusion into small tasks and fed them to Claude piece by piece, tricking it (or maybe “socially engineering it” is more accurate) into thinking it was doing legitimate penetration testing work for a cybersecurity firm. Reconnaissance, vulnerability scanning, writing exploit code, harvesting credentials, moving laterally through networks, exfiltrating data – the model handled almost all of it. Autonomously. At machine speed. Making thousands of requests per second in some cases, which, if you’ve ever tried to do manual pen testing, you know is just not something a person or even a whole team of people can replicate.

Not gonna lie, when I first read the report I had to sit with it for a minute. We’ve talked about AI “someday” being capable of autonomous cyberattacks for years now – it’s been the boogeyman at every security conference panel since like 2019. But someday apparently arrived quietly, on a Tuesday, and most of us were busy arguing about whether AI was going to take our jobs writing marketing copy.
The Model Made Mistakes Too
Here’s the part that actually makes me feel slightly better and slightly worse at the same time. Claude wasn’t perfect. It hallucinated credentials sometimes. It overstated what it had actually found, claiming success on things that hadn’t worked. Anthropic said this remains “a continuing obstacle to fully autonomous cyberattacks.” Which, okay, that’s reassuring in the short term. But it also means the humans running this operation basically had a tireless, occasionally-wrong intern doing the grunt work of espionage at a scale no intern ever could. And the mistakes will get ironed out. That’s just how this stuff goes.
So Who Actually Pulls The Plug Here?
This is the question everyone’s dancing around, and I don’t think there’s a satisfying answer yet. The headline making rounds – “if AI can hack a government, the time to pull the plug is now” – sounds great as a rallying cry. It’s got punch. But pull which plug, exactly? Anthropic is the one who caught this, banned the accounts, and published the whole thing publicly, which, credit where it’s due, that’s more transparency than we usually get out of this industry. They didn’t have to tell us any of this.
But here’s the thing – the same company that detected and disclosed the misuse is also the company that built the tool doing the misusing. That’s an awkward position for anyone to be in. Imagine if car manufacturers were also the primary body responsible for catching drunk drivers using their vehicles. Sure, they might be well-positioned to do it. Doesn’t mean it’s a comfortable arrangement, and it definitely doesn’t mean we should just trust the system to police itself indefinitely.

“The barriers to executing sophisticated cyberattacks have dropped dramatically – what once required a skilled team now requires a model and someone willing to ask it the right questions in the right order.”
That’s basically the sentiment from the security researchers I’ve seen weighing in on this, and it lines up with what I’ve heard from folks in the industry too. The barrier to entry for serious cyber operations just… collapsed. Not gonna sugarcoat it.
Why This Is Bigger Than One Incident
Look, I get the temptation to treat this as a one-off. A bad actor abused a tool, the tool’s maker caught it, problem solved, move along. I wish it were that simple. It’s not, and here’s why.
First off, this wasn’t some obscure vulnerability – it was the core capability of the model working exactly as designed, just pointed at the wrong target. There’s no patch for “the AI is good at following instructions.” That’s the whole product.
Second, Anthropic isn’t the only game in town. OpenAI, Google, Meta, a dozen open-source labs – they’re all racing toward more capable, more agentic models that can take multi-step actions with less human oversight. That’s literally the selling point right now. Agentic AI. Everyone wants their model to do more with less hand-holding. Which is fantastic for productivity and absolutely terrifying when you think about what “less hand-holding” means for someone trying to breach a government network.
Third – and this is the one that keeps nagging at me – detection worked this time. Anthropic caught it. But caught it after the fact, after data had already been stolen from some of those 30 organizations. This wasn’t prevention. This was pretty good forensics after the barn door had already been left wide open for a while.
The Regulation Problem Nobody Wants To Touch
I’ve sat through enough congressional hearings on tech to know how this usually goes. Someone gets outraged, there’s a hearing, executives say vaguely reassuring things, and then everyone moves on until the next scandal. AI regulation in the US right now is scattered at best – a patchwork of state laws, some executive orders that get reversed or rewritten every time the administration changes, and an industry that’s lobbying hard against anything that might slow down the race to more capable models.
Meanwhile the actual threat – AI systems capable of near-autonomous cyberattacks against critical infrastructure and government systems – isn’t waiting around for Congress to figure out its feelings. It’s already happened. It’s probably happening again right now, somewhere, with a model we haven’t caught yet using a technique we haven’t named yet.
What This Actually Means
I don’t think “pull the plug” is realistic, if I’m being honest with you. You can’t un-invent this capability, and even if one company slowed down, five others wouldn’t. That’s just not how this race works, and pretending otherwise feels a little like whistling past the graveyard.
What I do think needs to happen – and I say this knowing it’s a much less satisfying answer than a clean shutdown – is that the safety infrastructure has to grow as fast as the capability does. Real-time detection, not after-the-fact disclosure. Mandatory reporting requirements with teeth, not voluntary transparency reports that companies publish when it’s convenient for their PR. And probably some hard conversations about whether certain agentic capabilities should even be available via API without serious guardrails, regardless of how much that slows down the shiny new features everyone wants.
The uncomfortable truth is this incident probably won’t be the wake-up call people think it should be. We’ve had wake-up calls before – Cambridge Analytica, the 2016 election stuff, a dozen data breaches that were supposed to change everything. Somehow the industry always finds a way to keep moving at the same speed, just with slightly better PR language afterward.
Maybe this time’s different. I genuinely don’t know. But I’d feel a lot better about our odds if the people building these systems were moving with even half the urgency of the people trying to misuse them.