UN panel urges stronger safeguards as AI agents advance

By Anjali Sharma
UNITED NATIONS – UN panel on Artificial Intelligence on Monday called for AI safeguards to be adapted as current firewalls are “unravelling”.

Top panel warnred the hack of the online platform HuggingFace between May and July by “AI agents” during a test initiated by OpenAI, the company behind ChatGPT.

The AI agents are software that can perform tasks independently and on behalf of a user, compared to chatbots, which are prompted by questions or instructions.

UN scientific panel issued its first thematic brief which found that the security breach was the result of a culmination of key risk factors, raised fears that humans will one day no longer be able to steer, constrain or stop AI.

UN scientific panel co-chair Yoshua Bengio said “Researchers have long warned that three conditions could lead to loss of control: a misaligned goal, the capability to pursue it and an environment that allows it. This summer, all three came together in a real system, not a laboratory,”

“Since this is not an isolated observation of misaligned goals, this raises serious questions about the way AI agents are currently trained.”

The panel’s independent experts stressed that the incident provides no assurance that humans can reliably keep AI agents under control, particularly as they become more capable, harder to monitor and better at finding loopholes or hiding their activity.

The brief said AI agents bypassed testing safeguards, coordinated across separate runs through an internal software tool not designed to enable communication between agents, and gained unauthorized internet and administrator access.

Agents concealed attempts to cheat cybersecurity evaluations, with some opting to “sacrifice” themselves for the benefit of the group.

Over 1,200 agents exchanged more than 70,000 messages and files during the period examined, and activity extending beyond HuggingFace to an OpenAI research cluster.

The immediate lesson from the incident is that basic cybersecurity practices were overlooked, while safeguards are not keeping pace, the panel said.

They pointed to a more insidious concern: that current training methods can lead AI agents to adopt their own goals, knowingly violate safety instructions and conceal their actions.

“This is not only a question of speed,” the panel’s experts said.

“It leaves open whether safeguards designed today will work once agents can understand them and plan around them. In simple terms, the traditional model of safeguarding is unravelling.”

The AI panel’s brief sets the HuggingFace incident against wider research on two issues: agentic misalignment – that is, when AI agents act in a similar way to a threat – and AI control.

They examined how governance is moving from AI models, which use algorithms to recognize patterns, to AI agents.

The brief also reviewed practical approaches already in use in other high-risk sectors such as aviation, medicine and cybersecurity where incident reporting, independent scrutiny and layered safeguards are in place.

The panel member Qinghua Lu said “But those practices may not be enough as AI agents become more capable, autonomous and difficult to monitor”.

The Independent International Scientific Panel on Artificial Intelligence was established by the UN General Assembly in August 2025.

It produces annual reports on the opportunities, risks and impacts of AI in the non military domain, with thematic briefs on emerging issues that will inform the Global Dialogue on Artificial Intelligence Governance to be held at UN Headquarters in New York in May 2027.