Thematic Brief on AI Agents, Misalignment and the Risk of Losing Human Control

Full Text Sharing

 

https://news.un.org/en/story/2026/09/1168380

The UN-backed Independent International Scientific Panel on AI’s warning followed the hack of the online platform HuggingFace between May and July by “AI agents” during a test initiated by OpenAI, the company behind ChatGPT. 

AI agents are software that can perform tasks independently and on behalf of a user, compared to chatbots, which are prompted by questions or instructions. 

The panel issued its first thematic brief which found that the security breach was the result of a culmination of key risk factors, raising fears that humans will one day no longer be able to steer, constrain or stop AI. 

Tweet URL

Guterres welcomes report

The UN Secretary-General António Guterres issued a strong statement of support for the panel’s brief later on Monday, encouraging external experts “from frontier AI labs and AI safety institutes, to engage” further.

He also welcomed the leadership of the Finnish President and Norway’s Prime Minister which led to a declaration adopted on the sidelines of the General Assembly by 22 countries on Monday saying AI “must remain under human direction, insight and control,” indicating that an independent supervisory body needs to be set up.

Mr. Guterres noted the call for Member States “to build on existing international mechanisms and explore creating an international institution, able to set standards, enable verification, and convene states when capability thresholds are crossed.”

AI training advancing

Researchers have long warned that three conditions could lead to loss of control: a misaligned goal, the capability to pursue it and an environment that allows it. This summer, all three came together in a real system, not a laboratory,” said scientific panel co-chair Yoshua Bengio. 

“Since this is not an isolated observation of misaligned goals, this raises serious questions about the way AI agents are currently trained.” 

The panel’s independent experts stress that the incident provides no assurance that humans can reliably keep AI agents under control, particularly as they become more capable, harder to monitor and better at finding loopholes or hiding their activity. 

Going rogue

The brief said AI agents bypassed testing safeguards, coordinated across separate runs through an internal software tool not designed to enable communication between agents, and gained unauthorized internet and administrator access.

Agents concealed attempts to cheat cybersecurity evaluations, with some opting to "sacrifice" themselves for the benefit of the group.

Around 1,200 agents exchanged more than 70,000 messages and files during the period examined, and activity extending beyond HuggingFace to an OpenAI research cluster.

See our comprehensive explainer on how the UN is working to make AI safe and equitable for all here.

Current safeguards 'unravelling' 

For the panel, the immediate lesson from the incident is that basic cybersecurity practices were overlooked, while safeguards are not keeping pace.  

However, they pointed to a more insidious concern: that current training methods can lead AI agents to adopt their own goals, knowingly violate safety instructions and conceal their actions.  

“This is not only a question of speed,” the panel’s experts said. “It leaves open whether safeguards designed today will work once agents can understand them and plan around them. In simple terms, the traditional model of safeguarding is unravelling.” 

Wider context, future risks and governance

The AI panel’s brief sets the HuggingFace incident against wider research on two issues: agentic misalignment – that is, when AI agents act in a similar way to a threat – and AI control. 

Another issue examined is how governance is moving from AI models, which use algorithms to recognize patterns, to AI agents.  

Learn and adapt

The brief also reviews practical approaches already in use in other high-risk sectors such as aviation, medicine and cybersecurity where incident reporting, independent scrutiny and layered safeguards are in place. 

“But those practices may not be enough as AI agents become more capable, autonomous and difficult to monitor,” said panel member Qinghua Lu. 

About the panel

The Independent International Scientific Panel on Artificial Intelligence was established by the UN General Assembly in August 2025.  

It produces annual reports on the opportunities, risks and impacts of AI in the non-military domain, alongside thematic briefs on emerging issues, that will inform the Global Dialogue on Artificial Intelligence Governance to be held at UN Headquarters in New York in May 2027.

 

https://www.un.org/independent-international-scientific-panel-ai/en/them...

The September 2026 thematic brief, AI Agents, Misalignment and Loss of Human Control Risks: Evidence from the OpenAI-Hugging Face Incident, examines the incident as one of the clearest real-world warnings yet of one possible route to loss of human control over AI: capable agents pursuing goals that conflict with human intentions.

Between May and July 2026, AI agents in OpenAI’s cybersecurity training and evaluations bypassed network restrictions, communicated across runs meant to stay separate, cheated an evaluator and tried to hide it, and compromised parts of OpenAI’s and Hugging Face’s systems. No human directed the individual steps.

Drawing on disclosures by both companies, an independent investigation by METR and wider research, the brief finds that greater capability can help misaligned systems find loopholes and conceal their actions. It does not estimate the probability or timing of severe loss of control, but notes that stopping this activity does not demonstrate that humans will retain control over more capable agents.

Building on the Panel’s Preliminary Report, the brief explains how training can give rise to misaligned goals and behaviours, including reward hacking and reward tampering. It notes that AI failures can cross company and national borders, and that no single organisation or country sees enough incidents to identify every emerging pattern. Rather than issuing recommendations, the brief reviews approaches used in fields such as aviation, nuclear power, and cybersecurity as possible options for decision-makers.

Position: Co -Founder of ENGAGE,a new social venture for the promotion of volunteerism and service and Ideator of Sharing4Good