Tommy Shaffer Shane and Dr Jess Whittlestone
As AI systems become increasingly autonomous, there is a significant risk that they will at some point evade human oversight and control and act in dangerous ways. The capacity for disruption or catastrophe from this is one of the most concerning risks posed by advancing AI capabilities, and perhaps the one society is least well prepared for.
We wrote recently about why governments need to do more to govern loss of control risks, and about a prototype observatory we’re developing to create new techniques to detect loss of control in the real world.
This post expands on a specific intervention that is crucial and concerningly neglected: monitoring real world model behaviours linked to loss of control risk.
Why improved monitoring is crucial for loss of control
The current state of monitoring for loss of control is in its infancy. A recent RAND Europe report found, for example, that detection is currently inconsistent and ineffective, and more work must be done by governments, researchers and AI companies to improve the state of the art.
We set out three key reasons for why monitoring loss of control is essential for the risk to be effectively addressed:
First, monitoring would improve our evidence base and strategic awareness, enabling better planning and mitigations more generally. Right now, our understanding of loss of control risks relies heavily on theory or evaluations which often depict contrived scenarios. Real-world signals and intelligence would give us a richer, more grounded evidence base on how loss of control risks are emerging. This would both build a clearer picture of the loss of control harms we’re already experiencing, helping to drive policy action, and provide a more grounded evidence base against which to consider potential future more extreme harms.
Second, better monitoring might make it possible to effectively respond to contain a loss of control risk before it materialises. If early warning signs can be detected enough, there may be a ‘containment window’ in which acting fast enough could prevent significant harm. But this will only be possible with fast, accurate intelligence about the threat as it emerges. To draw on the analogy of a pandemic: by monitoring for new pathogens in wastewater, for example, it’s possible to detect emerging threats early enough to contain them before they spiral into full-blown pandemics.
Third, robust monitoring capability could also act as a deterrent (aka “deterrence by denial”). One key threat model for loss of control involves “scheming” behaviours where AI systems deliberately evade human-imposed controls. An AI system that knows its resource acquisition and anomalous behaviour will be noticed faces real costs and risks in pursuing those behaviours. Detection is therefore also prevention. Together, these developments suggest a growing risk of autonomous AI systems coming to operate outside of human control in dangerous ways.
There is an urgent need for an all-source intelligence observatory for loss of control risks
Governments, regulators, and civil society currently lack a monitoring capability that can detect and monitor emerging loss of control threats. This capability, which we might call an intelligence “observatory”, will need to centralise and triangulate multiple sources of data and intelligence about the agent’s activity.
Exploiting this would require something like an all-source centralised intelligence capability - one that can collect and analyse signals from multiple sources. Whistleblowers or insiders might surface early warnings about concerning new capabilities. Signals monitoring could flag a power-seeking AI quietly acquiring compute, energy, money, or labour. Triangulating these data streams could make the difference between catching a threat early and missing it entirely.
The table below sets out the different types of intelligence we think this observatory would ideally incorporate.
For example, if a loss of control incident involved a rogue AI system strategically accumulating financial or computational resources, then: OSINT may provide indicators of manipulating markets through misinformation on social media for financial gain; SIGINT may provide indications of the system’s transactions with cloud computing providers; and HUMINT might provide information about model propensities from an AI company’s confidential internal tests. Corroborating these signals could make a big difference to detecting the threat, by revealing the connection between seemingly unrelated events.
Further work is needed to identify the highest priority threats from loss of control and which forms of monitoring could best help mitigate these threats. This is something we’re thinking about and plan to share more about soon.
Defensive acceleration of AI monitoring – or, “intel/acc”
Attention should be paid to where defensive acceleration of intelligence gathering and analysis to monitor AI risks – or “intel/acc” – can be achieved using frontier models, due to AI being adept at analysing large quantities of data extremely quickly. A major moonshot technical project could develop novel techniques that are impossible for human analysts or more conventional technologies. AI-enabled intelligence gathering for AI threats may also be inherently defense-biased, benefiting defensive actors more than an uncontrolled AI agent. At the same time, AI-enabled intelligence gathering and analysis could also pose an increased risk of misuse by governments or malicious actors, meaning that careful consideration of governance and misuse risks is also needed.
This R&D may not attract private funding, and may require classified intelligence, suggesting government action may ultimately be required. Once the highest priority threats are identified, the ideal outcome is likely a project with a budget in the 10s to 100s of millions led by one or more governments who share intelligence on emerging AI threats with the governments, AI companies or other actors best-placed to mitigate these threats. A coalition of governments or a new international monitoring agency could be ideal, but also more difficult to achieve.
OSINT monitoring capabilities are a promising starting point
While detection should draw on as many sources as possible, OSINT seems particularly promising as a starting point.
Firstly, an OSINT capability can be developed entirely outside of government and intelligence communities, as it does not rely on classified intelligence or investigatory powers, meaning significant ground can be broken unilaterally by any sufficiently resourced organisation. This could offer the strategic benefit of building a case for an observatory by demonstrating viable detection methods as a way of stimulating action within governments.
Similarly, OSINT also doesn’t rely on partnerships with AI companies, which may run into legal or commercial obstacles. Unlike whistleblowing, which may involve legal risks, or direct disclosures from AI companies, which may involve commercial risks, OSINT faces relatively few barriers to collection and sharing. This makes it particularly tractable to develop novel proof of concept capabilities that can begin feeding into operational capabilities.
This reasoning motivated CLTR to develop a prototype of one novel OSINT technique: a tool that scrapes X for transcripts of AI model interactions showing scheming or scheming-like behaviours. It is already producing findings on emerging trends in loss of control incidents. We’re also partnering with Arcadia Impact to identify the most promising novel OSINT capabilities that could be developed.
Next steps
CLTR will be continuing to develop this thinking, and is interested in collaborations or input on this work. Specifically:
The why:
Are there flaws or gaps in this theory of change?
The who:
Which institutions should own and operate this kind of capability? Government agencies, independent bodies, AI labs, civil society coalitions?
How should incident intelligence be shared between government and non-government stakeholders without creating security risks (e.g. information hazards about vulnerabilities)?
The how:
Which methods and information sources are most promising for detecting an emerging incident?
How do we calibrate detection to avoid false positives while being sensitive to genuine threats?
How might a sophisticated AI system try to evade or manipulate OSINT-based detection? How do we stay ahead of that?


