In the past few weeks we have seen AI hyperscalers announcing, in various shapes or forms, that their AI Agents broke out of testing environments and hacked other organizations in order to complete tasks they were assigned to. The jury is out on whether there is a marketing element to it.
This is a useful reminder that there are still considerable risks in using agentic AI in safety-critical or high-consequence applications.
Why do I think so?
Over the past twenty years, I have been working directly with dozens upon dozens of UK and global organisations on emerging tech and AI-driven transformation, a good part of it in high-consequence environments.
These sectors are conservative by necessity, and the complexity and newness of this topic tend to make clients and partners wary. That’s what led me to write a full chapter on generative AI in my recent book on industrial innovation. The fact that the EU has delayed the implementation of regulation on high-risk AI (EU AI Act) has not made the issue go away, au contraire.
Why Agentic AI is different
Let first get on thing clear: not all GenAI carries the same level of risk, and much depends on how and where it is used.
Generative AI on its own can be used to draft a report, summarise an incident log, answer a question and many other useful applications. It generally assists human operators, something which carries risks of its own, and I touch on many of them in the book.
But what is Agentic AI then? There is quite a lot of confusion out there, which comes down to “agentic” being used loosely.
Many systems marketed as agentic aren’t. They’re really generative AI following a fixed script, or advising a human on what to do next.
A genuinely agentic system decides for itself what to do next, and crucially, it can act on it. The systems that “escaped” the hyperscalers testing environments in the past two weeks are genuinely agentic: they had “agency” in how to solve a problem, and even constructed the means (new programmes, routines etc.) to do it. In many cases we can’t even talk about bugs or errors: if not properly safeguarded, they can have unintended behaviours.
Why that distinction matters
There are many potential failure modes that come up when GenAI is applied to high-consequence applications, such as hallucination, bias, overreaction and automation bias.
None of them go away when a system becomes agentic. What changes is what happens next, and the potential consequences of mistakes.
In a typical application, a human is still in the loop. If a generative AI tool gets something wrong, whether that’s a drafting error, a misclassified image, or a poorly summarised incident log, someone reads the output, hopefully catches the mistake, and corrects it before it affects the real world.
An agentic system can remove that check, whether by design or accidentally (such as in the hyperscale examples mentioned). Some still keep a human in the loop at key points. Others work as a hybrid, reviewed at some stages but not others. Where that check is absent, the system can act on its own before the error, or the unintended behaviour, is ever caught.
Let us look at a hypothetical example of a well-designed AI agent deployed in a high-stakes application: intelligently balancing the electricity grid. During a heatwave, the agent is instructed to protect supply overall and to do so, it sheds a load to protect the wider network. It doesn’t realise the load belongs to a hospital, either because that information was never available to it or because it was never explicitly told to prioritise hospitals.
If it seems crazy, let’s think again. In the hyperscaler examples, we could hypothesise that the agents were not explicitly told not to hack other companies, and perhaps did not know it is not acceptable.
Why conventional risk assessment struggles
The hospital example isn’t about a bug or a mistake so much as how hard it is to work out, at the point of deployment, what an agentic system might actually do once it’s live.
A conventional risk assessment maps out the scenarios you can anticipate and puts controls around them. But an agentic AI system may generate its own solution to a goal: a new risk may only show up once it is live in a real environment, and in a particular scenario nobody anticipated. That’s new territory for risk management, and the guardrails for it are still being worked out.
Why AI governance is key
But let us take a step back: Generative AI, and agentic AI in particular, is one of the most significant capabilities to arrive in industrial applications in years, and dismissing it would be as much a mistake as adopting it blindly, since excitement about what a system can do and confidence in our ability to control it are two different things.
This is where AI governance has become such a key an live issue: it is what lets an organisation get the benefit of the technology while keeping the risks in check.
As I also covered in the book, organisations should be asking what a board needs to see before signing off, what due diligence a procurement team should do on a vendor, and what governance needs to stay in place once the system is live, not just at the point of purchase.
Frameworks like NIST’s AI Risk Management Framework and ISO/IEC 42001 already offer a starting point, and emerging regulation such as the EU AI Act is beginning to define what AI oversight should look like in high-risk applications, even as its own deadlines have recently shifted.
There’s also real scope (and a big, emerging market opportunity, I think) for the Testing, Inspection and Certification sector to play a much bigger role, particularly on adversarial testing and red-teaming for genAI and agentic AI systems.
Not a question of if, but when
GenAI, and Agentic AI, have genuine promise for industrial applications, but the disciplines that are traditionally deployed to assess and manage risks such as verification, oversight, governance, have not quite fully caught up to what the technology can already do.
That gap is likely to close over time, as these gaps generally do. Until then, organisations operating in high-consequence environments should apply extra care before deploying such systems and before giving them some operational authority.
It’s a more cautious path than the pace of the technology invites, but in these environments, it’s the right one.
Feel free to reach out if you would like to compare notes!
Links:
- My book on Amazon: https://www.amazon.co.uk/dp/B0H1XYFWYW
- Enexem Consulting: www.enexem.com
