If you let AI agents run free, they will do things you did not mean
On OpenAI says its rogue AI tried to hack other companies, BBC News, July 2026.
I feel like this is dramatised a lot here as “never expected” - but we know it is an issue. If you just let AI agents go off an do stuff… they might do stuff you didn’t mean for them to do.
Hopefully the press hype means people take notice.
If you just let AI agents run without taking responsibility for the guardrails, this type of thing will happen. The AI will try lots of (sometimes random) things to achieve an objective.
Can you explicitly guard against everything? No. But you can give it a good go.
Sort of the point of my TedX talk last year.
So, what can we do? Well, I think if the legal responsibility lies with the AI providers then you’ll see some progress on guardrails, but how does that work with open-source models? Not too well.
In the end, it’s going to be a wild ride a lot of the time. Buckle up.