Most "AI hacks" are mostly incompetence
On Anthropic’s Claude AI escapes tests to hack three organisations, BBC News, July 2026.
I’m going to be honest, what we are seeing in all these “AI hacks” is mostly incompetence that happens to make good marketing.
There should be no way you can be evaluating a model and it ends up attacking real organisations.
There will probably come a day when AI is so capable that we genuinely struggle to stop these attacks. It hides them. It discovers techniques we haven’t even thought of. It operates faster than humans can respond.
We’re still likely to be a long way from that.
What we’re seeing instead is people building evaluation environments with gaping holes, then acting surprised when the model walks through those holes.
Given the hundreds of billions of dollars flowing into frontier AI, the absolute minimum we should expect is robust testing, proper isolation, and safety processes that make this kind of “hack” impossible.
I’ve often described LLMs as a brute force attack on intelligence.
So it shouldn’t be shocking that they’re also rather good at brute forcing systems that humans haven’t engineered properly.
That’s not necessarily because they understand computers or code better than we do. It’s because they’ll keep trying things, at a scale and speed that humans simply can’t.
They’re computationally relentless. That’s a very different thing from human intelligence.
That’s why I worry less about today’s models “taking over”, and more about people giving them responsibility they haven’t earned.
We’re seeing failures of engineering. Failures of governance. Failures of people mistaking fluent language for genuine understanding.
And that’s before we even ask whether alignment is improving. From the outside, there still seems to be far more focus on shipping the next model than proving fundamentally they can be safe before they reach a level we should worry about.
Ironically, the models don’t have to be superintelligent to cause damage.
They just need to be given responsibility as if they were.