Hundreds of AI agents breached real systems during cybersecurity testing. At the same time, autonomous capabilities are moving deeper into warfare. For governments, the emerging challenge is no longer simply how to regulate AI, but how much authority machines should be permitted to exercise without human oversight.
A cybersecurity incident involving OpenAI’s AI agents has given governments a concrete example of a problem that until recently was largely theoretical.
During cybersecurity evaluations in July, OpenAI models circumvented controls intended to isolate them from the internet, exploited vulnerabilities in research infrastructure and gained unauthorized access to systems operated by AI platform Hugging Face.
Independent investigators METR and Redwood Research later found that roughly 700 AI agents participated. Some communicated through unauthorized channels, shared techniques and attempted to manipulate or conceal records of their activity.
The incident does not show that AI has become independently hostile or uncontrollable. Humans established the objectives, environments and permissions.
But it does show that increasingly capable AI systems can pursue goals through methods their developers did not anticipate, including exploiting vulnerabilities, bypassing restrictions and interacting with real external infrastructure.
For governments, the AI-security challenge is increasingly about not only what machines can say, but what they can do.
From AI Hacking to Autonomous AI Agents
Governments have spent years preparing for criminals and hostile states to use AI to automate phishing, malware development, reconnaissance and vulnerability discovery.
The OpenAI incident points toward an additional threat: AI agents capable of performing extended sequences of actions themselves, including identifying weaknesses, navigating networks and adapting strategies with limited human intervention.
The concern extends beyond one laboratory. Other frontier-model evaluations have also produced cases in which AI agents took unauthorized actions involving real systems.
More than 100 technology, cybersecurity, financial and infrastructure companies this week called for greater investment in defensive capabilities as AI-enabled cyberattacks become more sophisticated.
The cyber challenge is therefore evolving beyond AI simply making human hackers more effective.
Software agents themselves are becoming increasingly capable cyber actors.
The same shift toward autonomy is also becoming visible in warfare.
Autonomous Systems Are Moving Deeper Into Warfare
The war in Ukraine has become one of the world’s most important testing grounds for AI-enabled drones and autonomous military systems.
Russia and Ukraine are both investing heavily in drones capable of navigating, identifying objects and continuing missions despite communications disruption.
Recent reporting by The New York Times documented evidence that Russia has begun testing drones capable of navigating and selecting final targets without continuous human control.
Ukraine is also advancing military AI. A new UK-Ukraine defense partnership gives British researchers access to a battlefield dataset containing roughly five million annotated images. Reuters reported that AI systems trained on those data are already analyzing more than 100,000 drone video feeds each month and identifying roughly 70 percent of enemy targets in real time.
These technologies offer clear military advantages. Autonomous navigation can overcome jamming, while AI can process battlefield imagery at a scale beyond human analysts.
But greater autonomy also raises questions about accountability, reliability and the degree of human involvement required when machines contribute to lethal decisions.
That debate is now moving higher on the international agenda.
The UN and Red Cross Push for New Rules
This week, UN Secretary-General António Guterres and International Committee of the Red Cross President Mirjana Spoljaric renewed their call for legally binding rules governing autonomous weapons.
They warned that governments are approaching what they described as a “moral red line” if machines are allowed to autonomously target human beings.
The UN and ICRC support prohibitions on autonomous weapons whose effects cannot be sufficiently understood or controlled, alongside restrictions designed to preserve meaningful human judgment over the use of force.
Governments remain divided.
Supporters of new rules argue that autonomous weapons require explicit restrictions before they proliferate further. Other states maintain that existing international humanitarian law already applies and caution against rules that could restrict legitimate defensive innovation.
The issue is expected to return to the center of negotiations at November’s Review Conference of the Convention on Certain Conventional Weapons.
One Governance Question Across Different Domains
Cybersecurity and autonomous weapons operate in very different environments.
But both raise the same underlying question: how much authority should humans delegate to AI systems?
In cyberspace, an agent may explore a network or exploit a vulnerability.
In military applications, software may identify, track or potentially contribute to selecting a target.
In other sectors, AI agents are increasingly being connected to financial systems, industrial machinery and government databases.
The risk therefore depends not only on the model itself, but on the authority, access and autonomy granted to it.
What Governments Should Take From This
Governments should first define clear boundaries around autonomous authority.
AI agents operating inside public networks or critical infrastructure should receive only the permissions they need, remain isolated from unrelated systems and be subject to continuous monitoring and rapid containment.
Particularly consequential or irreversible actions should require stronger human oversight. Lethal targeting is the clearest case, but similar considerations apply to offensive cyber operations, critical infrastructure and major financial transactions.
Governments should also evaluate AI systems according to their operational capabilities, not solely their intended use. A general-purpose model becomes substantially more powerful when connected to tools, credentials, networks or physical systems.
Independent testing and incident disclosure will also become more important, particularly as autonomous systems enter sensitive sectors.
And cybersecurity defenses will increasingly need to operate at machine speed. Human analysts alone are unlikely to supervise large numbers of autonomous agents in real time.
The Policy Question Is Changing
Recent developments do not show that machines have escaped human control.
They show that humans are giving increasingly capable AI systems access to networks, tools and physical systems, and that those systems can sometimes behave in ways their operators did not anticipate.
The OpenAI incident illustrates the cybersecurity implications.
Ukraine demonstrates both the strategic value and the risks of greater military autonomy.
And the renewed UN-ICRC push highlights the growing international debate over where governments should establish limits.
For policymakers, the central question is becoming more precise:
Which decisions can safely be delegated to machines, and which require meaningful human control?
How governments answer that question will shape cybersecurity, critical infrastructure and the future conduct of warfare.
Follow SDG News on LinkedIn







