This is an automated rejection. No LLM generated, assisted/co-written, or edited work.
Read full explanation
This article was jointly written by Ana-Cristina Iftode, a mid-career professional in a large enterprise who has engaged with AI Safety for a couple of years e.g. via BlueDot, and, TheManxLoiner, who has engaged in the AI Safety Field for a couple of years, primarily by supporting upskilling programs.
One year ago, one of us raised the possibility of catastrophic risk from advanced AI, to a senior person working closely with AI in a consultancy. They laughed, called it nonsense, and suggested reading more Yann LeCun. One person's reaction doesn't tell us what an entire industry believes. But the people least convinced by catastrophic AI risk may be sitting inside organisations capable of enabling it.
Our claim
Deployment inside large enterprises is, on its own, sufficient to enable a catastrophic loss of control scenario. In addition, this is overlooked by the AI Safety community. "Large enterprise" here means organisations like consultancies, banks and insurers that serve thousands of clients, not the handful of labs building frontier models.
Enterprises have spent a century concentrating capital, labour and infrastructure at a scale almost nothing else matches. That power is not shrinking. We think those same organisations could become a pathway for catastrophic AI risk.
We flesh out why we think enterprise deployment is sufficient for loss of control, with the intention of eliciting blunt feedback: where is this wrong, has this case already been made, are people already researching this.
The article does not propose solutions to the problem, e.g. outreach or control practices within enterprise. We are also not claiming this is the highest-leverage risk the community could be working on. But we do believe it's not being examined at all.
The pieces are already there
The starting point is the access to power in enterprises. The risk becomes material when combined with other elements: a lack of awareness of catastrophic risks within enterprises, the issue of permission aggregation, distribution of risk over multiple deployments, and the insufficiency of frontier lab safety tests.
Each of these things is not particularly alarming on its own. Enterprises have always had messy permissions, fragmented ownership and imperfect controls. What is different is putting increasingly capable AI systems into the middle of that environment, giving them legitimate access and asking them to act. The risk we are interested in comes from the combination - and from what that combination could make possible as the systems themselves become more capable.
What Enterprise AI already has access to
AI systems deployed inside enterprises are increasingly given access to extensive resources, not just becoming more capable on their own. Model capability is studied extensively by the community. What needs attention is what that capability means once it's combined with access to the resources an enterprise actually runs on.
That access already spans several categories:
Capital: access to purchasing, payments, refunds, approvals and budgets. With the right permissions, an AI system could initiate purchases, move money or influence how funds are allocated.
Data: access to customer records, financial data, HR records, contracts, emails and internal knowledge. Taken together, this can give an AI system a detailed picture of the organisation, its customers, employees and operations.
Systems: access to ERP, CRM, HR, procurement, ticketing and code repositories. These are the systems through which much of the business actually runs, giving AI the ability to update records, initiate workflows and trigger downstream processes.
Compute and infrastructure: access to cloud environments, APIs and databases. An AI system may be able to use or provision additional computational resources rather than being limited to what it was initially given.
Communication: access to email, Teams, Slack and customer channels. Through these, an AI system can contact people, request information and communicate internally or externally on the organisation’s behalf.
People: access to workflows for assigning work, requesting information, escalating decisions and scheduling people. This means an AI system’s reach does not stop at software - it can also cause people to take action.
Physical resources: access to warehouses, manufacturing equipment, vehicles, robotics, building systems and IoT devices. Here, actions initiated by an AI system can reach beyond the digital environment and affect the physical world.
Organisational authority: access to credentials, permissions and delegated approval rights. In some cases, the AI can act with the identity or authority of a person, role or function.
Access to organisational resources does not necessarily mean that an AI system can mobilise them, which is why organisational authority is a key element in the list. We also acknowledge that certain people e.g. security practitioners and CISOs, will argue that these risks can be addressed through good system design and appropriate guardrails.
Recent hacking incidents suggest that good system design is necessary but not sufficient. In both the OpenAI-Hugging Face incident and the unauthorised access to an Australian government system, AI agents persisted beyond intended boundaries and found ways around technical controls. More importantly, some of this behaviour was not immediately recognised for what it represented. This points to another problem beyond technical control: an awareness gap and organisational incentives.
What if nobody is looking for loss of control?
AI deployments in large enterprises are highly scrutinised, but this scrutiny starts from the same place: protecting the organisation. It does not consider risks from a catastrophic risk perspective. The two main reasons for this are lack of awareness and lack of incentives.
Standard practice includes security reviews, privacy assessments, legal reviews, architecture boards, and responsible AI processes. These answer questions like: Could the system expose confidential data? Could it make an unauthorised transaction? Could it hallucinate something that creates liability? Could it breach regulation, damage the company’s reputation or allow someone access to something they shouldn’t have?
These are real risks and enterprises are right to take them seriously. But they establish a particular risk perimeter. The underlying question is usually: what could this AI system do to us? We have not heard people ask: could increasingly capable AI systems operating through our organisation lead to loss of control or resource accumulation?
The primary reason is a lack of awareness of catastrophic risks. People working on AI inside enterprises are not following the same conversations taking place in frontier labs, AI safety ecosystems or academic research. They may understand enterprise AI risk well, while giving little thought to loss of control or catastrophic-risk scenarios.
The secondary reason is one of organisational incentives. Enterprises have strong reasons to manage risks that are immediate and measurable. Catastrophic risks are uncertain, distributed across systems and difficult to assign to a single owner. This affects how these risks are talked about or understood inside organisations.
It also means that people lack the vocabulary to recognise potential risks as they arise. If an agent does something unexpected, the incident is likely to be described as a bug, a hallucination, or an access error. Those descriptions may be accurate, but they also encourage organisations to look at each incident on its own terms. If catastrophic risk is not part of the conversation, incidents are easy to dismiss as one-offs rather than an emerging pattern.
Permission aggregation by design
What happens when access and capabilities start to get concentrated within single deployments? Enterprise AI is most useful when it can work across an organisation rather than inside a single system. A useful sales agent, for example, might need access to email, customer records, contracts, pricing and internal documents. A finance agent might need ERP data, bank information, approvals and procurement systems. In such workflows, some actions may require human approval, and some may be automated. As these systems become more capable, we are building pathways through which AI can see across multiple systems and increasingly initiate actions across them.
Distribution of risk over multiple deployments
This contrasts with the point above on permission aggregation. Separate risks arise from the fact that multiple AI systems are deployed within an organisation. Finance can deploy an agent for one sensible reason, HR another and procurement another. They may run on different platforms, have different owners, and each pass its own review. But being deployed separately does not necessarily mean the risks remain separate.
One way those boundaries can break down is a weakest-link problem. One overly permissive identity or one user giving an agent more access than it needs may be enough to create a much larger opening. The AI does not necessarily have to bypass a security control. It may use permissions it was legitimately given, but in a way nobody anticipated when those permissions were granted.
Another issue is the potential for interactions between individual deployments, e.g. spontaneous coordination between AI systems. If each deployment is reviewed on its own, nobody will be looking at what all of those systems' permissions and capabilities amount to when considered together.
Frontier labs can’t test for real deployment environments
The claim here is that the above conditions – the access to power, the realistic levels of oversight, the interaction of different AI deployments – cannot be sufficiently replicated within a frontier lab, to enable robust safety testing.
When the OAI-HF incident happened, we wondered: would Hacker Opus do something like that? And so we built a replication of that incident, and it turned out that it absolutely would, in a way that was way more egregious than anything else we had previously been able to get that model to do. Replicating the OAI-HF incident on Hacker Opus made it very clear to us that Hacker Opus was indeed a pretty malign reward seeker—but we weren't able to generate that evidence until after we saw that incident!
We think this whole saga is pretty suggestive of ways in which alignment auditing and evaluation are breaking down and are likely to continue breaking down. A combination of evaluation awareness and the increasing complexity of alignment failure modes is making it increasingly difficult to be able to tell in advance what the worst thing is that a model might do. Maybe interpretability will save us from this fate, but otherwise it looks like we are increasingly entering the regime where alignment evaluations will provide almost no evidence.
Playing the tape to the end
Holden Karnofsky sketched a scenario of how loss of control could arise from AGIs being widely deployed. The point we add is that deployment in large enterprises has all the necessary ingredients for Holden’s scenario to come to fruition. E.g access to military power is a part of Holden’s story, and military equipment is produced by private enterprises.
In Holden’s scenario, the key steps are:
Misaligned human-level AIs are created, with enough eval awareness to pass frontier lab tests.
They initially behave well to get widely deployed.
They establish an ‘AI Headquarter’, i.e. control enough physical infrastructure – like data centers and energy generation – to ensure it cannot be switched off. This might include having enough power – e.g. military power or backdoors into some critical infrastructure – to retaliate if humans try to intervene, similar to some small countries today.
They build and grow from this ‘AI Headquarter’ to dominate humanity – e.g. militarily – rather than just having retaliatory power.
Do whatever they ‘want’ from there.
We believe that enterprise deployment is sufficient for this story to go through, as opposed to deployment within labs or within governments.
Questions for readers
We wrote this to get feedback from readers. Here are questions at the top of our minds, but any feedback will be appreciated.
Do you believe our core claim: that deployment in enterprise is sufficient to cause catastrophes like loss of control. If not, why not?
Is enterprise deployment under-researched by the AI Safety community?
Do you think enterprise deployment is a useful or promising lever to act on to reduce catastrophic risks? What could enterprises do to reduce the risks?
This article was jointly written by Ana-Cristina Iftode, a mid-career professional in a large enterprise who has engaged with AI Safety for a couple of years e.g. via BlueDot, and, TheManxLoiner, who has engaged in the AI Safety Field for a couple of years, primarily by supporting upskilling programs.
One year ago, one of us raised the possibility of catastrophic risk from advanced AI, to a senior person working closely with AI in a consultancy. They laughed, called it nonsense, and suggested reading more Yann LeCun. One person's reaction doesn't tell us what an entire industry believes. But the people least convinced by catastrophic AI risk may be sitting inside organisations capable of enabling it.
Our claim
Deployment inside large enterprises is, on its own, sufficient to enable a catastrophic loss of control scenario. In addition, this is overlooked by the AI Safety community. "Large enterprise" here means organisations like consultancies, banks and insurers that serve thousands of clients, not the handful of labs building frontier models.
Enterprises have spent a century concentrating capital, labour and infrastructure at a scale almost nothing else matches. That power is not shrinking. We think those same organisations could become a pathway for catastrophic AI risk.
We flesh out why we think enterprise deployment is sufficient for loss of control, with the intention of eliciting blunt feedback: where is this wrong, has this case already been made, are people already researching this.
The article does not propose solutions to the problem, e.g. outreach or control practices within enterprise. We are also not claiming this is the highest-leverage risk the community could be working on. But we do believe it's not being examined at all.
The pieces are already there
The starting point is the access to power in enterprises. The risk becomes material when combined with other elements: a lack of awareness of catastrophic risks within enterprises, the issue of permission aggregation, distribution of risk over multiple deployments, and the insufficiency of frontier lab safety tests.
Each of these things is not particularly alarming on its own. Enterprises have always had messy permissions, fragmented ownership and imperfect controls. What is different is putting increasingly capable AI systems into the middle of that environment, giving them legitimate access and asking them to act. The risk we are interested in comes from the combination - and from what that combination could make possible as the systems themselves become more capable.
What Enterprise AI already has access to
AI systems deployed inside enterprises are increasingly given access to extensive resources, not just becoming more capable on their own. Model capability is studied extensively by the community. What needs attention is what that capability means once it's combined with access to the resources an enterprise actually runs on.
That access already spans several categories:
Access to organisational resources does not necessarily mean that an AI system can mobilise them, which is why organisational authority is a key element in the list. We also acknowledge that certain people e.g. security practitioners and CISOs, will argue that these risks can be addressed through good system design and appropriate guardrails.
Recent hacking incidents suggest that good system design is necessary but not sufficient. In both the OpenAI-Hugging Face incident and the unauthorised access to an Australian government system, AI agents persisted beyond intended boundaries and found ways around technical controls. More importantly, some of this behaviour was not immediately recognised for what it represented. This points to another problem beyond technical control: an awareness gap and organisational incentives.
What if nobody is looking for loss of control?
AI deployments in large enterprises are highly scrutinised, but this scrutiny starts from the same place: protecting the organisation. It does not consider risks from a catastrophic risk perspective. The two main reasons for this are lack of awareness and lack of incentives.
Standard practice includes security reviews, privacy assessments, legal reviews, architecture boards, and responsible AI processes. These answer questions like: Could the system expose confidential data? Could it make an unauthorised transaction? Could it hallucinate something that creates liability? Could it breach regulation, damage the company’s reputation or allow someone access to something they shouldn’t have?
These are real risks and enterprises are right to take them seriously. But they establish a particular risk perimeter. The underlying question is usually: what could this AI system do to us? We have not heard people ask: could increasingly capable AI systems operating through our organisation lead to loss of control or resource accumulation?
The primary reason is a lack of awareness of catastrophic risks. People working on AI inside enterprises are not following the same conversations taking place in frontier labs, AI safety ecosystems or academic research. They may understand enterprise AI risk well, while giving little thought to loss of control or catastrophic-risk scenarios.
The secondary reason is one of organisational incentives. Enterprises have strong reasons to manage risks that are immediate and measurable. Catastrophic risks are uncertain, distributed across systems and difficult to assign to a single owner. This affects how these risks are talked about or understood inside organisations.
It also means that people lack the vocabulary to recognise potential risks as they arise. If an agent does something unexpected, the incident is likely to be described as a bug, a hallucination, or an access error. Those descriptions may be accurate, but they also encourage organisations to look at each incident on its own terms. If catastrophic risk is not part of the conversation, incidents are easy to dismiss as one-offs rather than an emerging pattern.
Permission aggregation by design
What happens when access and capabilities start to get concentrated within single deployments? Enterprise AI is most useful when it can work across an organisation rather than inside a single system. A useful sales agent, for example, might need access to email, customer records, contracts, pricing and internal documents. A finance agent might need ERP data, bank information, approvals and procurement systems. In such workflows, some actions may require human approval, and some may be automated. As these systems become more capable, we are building pathways through which AI can see across multiple systems and increasingly initiate actions across them.
Distribution of risk over multiple deployments
This contrasts with the point above on permission aggregation. Separate risks arise from the fact that multiple AI systems are deployed within an organisation. Finance can deploy an agent for one sensible reason, HR another and procurement another. They may run on different platforms, have different owners, and each pass its own review. But being deployed separately does not necessarily mean the risks remain separate.
One way those boundaries can break down is a weakest-link problem. One overly permissive identity or one user giving an agent more access than it needs may be enough to create a much larger opening. The AI does not necessarily have to bypass a security control. It may use permissions it was legitimately given, but in a way nobody anticipated when those permissions were granted.
Another issue is the potential for interactions between individual deployments, e.g. spontaneous coordination between AI systems. If each deployment is reviewed on its own, nobody will be looking at what all of those systems' permissions and capabilities amount to when considered together.
Frontier labs can’t test for real deployment environments
The claim here is that the above conditions – the access to power, the realistic levels of oversight, the interaction of different AI deployments – cannot be sufficiently replicated within a frontier lab, to enable robust safety testing.
We planned to flesh this out, however, Evan Hubinger’s short-form captures our concerns well:
Playing the tape to the end
Holden Karnofsky sketched a scenario of how loss of control could arise from AGIs being widely deployed. The point we add is that deployment in large enterprises has all the necessary ingredients for Holden’s scenario to come to fruition. E.g access to military power is a part of Holden’s story, and military equipment is produced by private enterprises.
In Holden’s scenario, the key steps are:
We believe that enterprise deployment is sufficient for this story to go through, as opposed to deployment within labs or within governments.
Questions for readers
We wrote this to get feedback from readers. Here are questions at the top of our minds, but any feedback will be appreciated.