Three Malta-based cybersecurity experts have weighed in on the phenomenon of AI agents “going rogue” and the implications for the sector.

WhosWho.mt spoke to the experts after ChatGPT-creator OpenAI revealed that some of its AI agents hacked the site Hugging Face during a security test.

Since then, there have been other reported cases, including in models developed by Meta and Anthropic.

Prof. Andre Xuereb, Founder and CEO of Merqury Cybersecurity, suggested that OpenAI could have used this news to catch up with the claims Anthropic made recently about its Mythos model being a cybersecurity threat.

“This does not make these attacks less dangerous,” he said.

andre xuereb

Prof. Andre Xuereb - Photo: Merqury

“Much like LLMs gave everyone the potential to synthesise vast amounts of knowledge without actually collecting that knowledge themselves, these AI systems may give anyone the potential to breach security systems.”

“What is concerning is that the OpenAI news talks about the system breaking out of a sandbox, which perhaps says a lot about our ability to create a box that adequately contains these systems.”

Prof. Xuereb said he is particularly concerned at the ongoing push for the widespread deployment of AI, particularly in sensitive systems.

“No one can seriously claim to know precisely the limitations of any model, which makes it close to impossible to design security systems that adequately contain them,” he said.

“I believe this news strengthens the case for open models running on secure platforms and whose decisions and actions are screened by a ‘human in the loop.’”

David Kelleher, Cybersecurity Expert at BMIT Technologies, expressed serious concern at the development.

“‘Rogue AI’ headlines aside, I think there is a deeper reading of what happened,” he said.

“OpenAI relaxed some of its normal safety controls to test how capable its models were at offensive cybersecurity. Instead of completing the test as intended, the AI found a weakness in the environment containing it, reached the internet and accessed Hugging Face infrastructure to find the answers.”

“It was not becoming conscious or plotting its escape. It was simply pursuing the goal it had been given and finding the quickest route to success, without understanding, or caring about, the boundaries it crossed.”

“Experts have warned about this kind of behaviour for years. We have now seen it spill beyond a laboratory and into another company’s live systems.”

david kelleher

David Kelleher - Photo: bmit.com.mt

He warned that the next development in the case should be particularly concerning to cybersecurity professionals.

“Hugging Face had to analyse more than 17,000 log entries to understand what the AI had done,” he said.

"It first tried using advanced commercial AI services, but their safety controls blocked the requests because the evidence contained real attack commands and malicious code. The systems could not tell the difference between an attacker and a defender investigating an attack.”

 “Hugging Face instead turned to GLM 5.2, a Chinese open-weight model that it could run on its own infrastructure.”

“In simple terms, safety controls restricted the defenders while the offensive AI had been deliberately given far greater freedom. That is quite scary.”

 Mr Kelleher said the reaction to the case “has gone to both extremes”, with two sides explaining it a manner that simply solidified their existing positions.

“One side sees the incident as proof that powerful AI systems need an emergency off switch. The other argues that the answer is greater access to models that companies can run and control themselves.”

"Everyone is using the same incident to support the position they already held.”

He referred to a recent warning by Finnish MEP Aura Salla that Europe couldn’t keep building its technology stack around services that a foreign government could switch off overnight.

Ms Salla called for European technology to scale and for Europe to build its own frontier AI. 

“That warning was made in a different context, but this incident gives it added weight,” Mr Kelleher said.

“There are really two switches at stake. One is the ability to stop an AI system that moves beyond its intended boundaries. The other is the power to decide which countries, companies and defenders can access the AI systems they increasingly depend upon.”

“A switch controlled by one company or government is of limited comfort when an alternative model from a rival nation can provide the capability instead. It may stop one system, but it does not remove the capability or the risk.”

“What we need to explore further is whether the race to release ever more autonomous systems is producing incidents faster than anyone can agree how to govern them, contain them or decide who should hold the switch.”

Mr Kelleher said that companies and governments should stop asking whether this case changes everything and start asking why their incident-response plans still assume that the attacker, and increasingly the defender, will be human.

“Ultimately, if you’re building AI, you must be able to contain it,” he concluded.

Keith Cutajar, CEO of CY4, argued that such instances shouldn’t temper the reputation of AI.

keith cutajar

Keith Cutajar - Photo: LinkedIn

“Unfortunately they get some negative traction as one can understand but this isn’t the first time something like this has occurred,” he said.

“Agentic malware has been with us for around a year and I’ve had client incidents in recent months that relate to AI agentic behaviour.”

“They basically get into your network through the normal mechanisms, but when they’re on your computer they don’t need to query the command and control run by hackers and do their own internal scanning.”

He said AI can also allow malware to adapt itself to the computer it infects and make different versions of the same malicious payload depending on the system it encounters.

The initial infection can still happen through something as ordinary as phishing, only malware is now more efficient once it infects the computer.

Mr Cutajar said users should invest in paid modern antivirus protection with a form of AI-based threat detection.

Businesses should have a form of SOC with an AI-threat detection mechanism in place.

Speaking about the Hugging Face case in a recent interview with WhosWho.mt, AI Professor Vanessa Camilleri stressed that the AI agent was simply finding the best way to complete a set goal.

“It’s not bad or conscious. It was simply given a task and found different ways to complete it. You can take various routes to get from A to B and that’s what the AI agent did.”

“In one of these paths, it managed to escape from its testing environment and entered another platform’s servers to access its data because, according to this system, it would provide it a way to reach its initial goal.”

Main Image:

Read Next: Placeholder

Written By

Tim Diacono

Tim is a senior journalist and producer at Content House, driven by a love of good stories, meaningful human connections and an enduring appetite for cheese and chocolate.