Governing AI Where It Acts: From Model Choke Points to Organizational Control Points

With Nalin Kulatilaka (Boston University)

Appeared on Substack on August 27, 2026. https://anandanandalingam649613.substack.com/p/governing-ai-where-it-acts?r=o7w77&utm_campaign=post&utm_medium=web&showWelcomeOnShare=true

Artificial intelligence is evolving faster than many regulatory and organizational processes. Yet, we believe institutions still devote too little attention to the practical question: how can they prevent or contain harm when AI is placed inside real workflows, connected to sensitive data and tools, and allowed to influence consequential decisions?

This is not an argument for deregulating AI. Nor is it an argument for treating every general-purpose language model as a regulated product merely because it is powerful. Development and ordinary access should remain relatively open, except where a model presents genuinely dangerous or systemic capabilities. Government has a critical role in setting and enforcing duties for consequential applications, often through sector regulators and existing safety regimes. Model-level safeguards remain necessary in exceptional cases, but they cannot substitute for controls where organizations actually deploy AI and give its outputs authority.

[AI risk also includes deliberate misuse by insiders or outside attackers, societal externalities such as energy demand or effects on children and work, and global threats from rogue or terrorist actors. Those risks may require stronger upstream controls and broader public coordination. They matter, but they are not our focus here. This essay concentrates on organizational failure: harm that arises when otherwise legitimate institutions use AI through weak objectives, flawed data, inadequate testing, poorly bounded permissions, diffuse responsibility, or ineffective monitoring.]

Consider a pharmaceutical company developing an antiviral for an emerging pathogen. It might use the same underlying AI system to review scientific literature, identify a biological target, rank possible drug molecules, interpret laboratory results, and monitor a clinical trial. It would make little sense to impose the same restrictions at every stage. Using a model to summarize published research is not equivalent to allowing its recommendation to initiate a laboratory experiment, select a clinical candidate, or influence the dosing of a patient.

The relevant question is therefore not simply who can use the model. It is when the model’s output acquires access to sensitive resources, institutional authority, physical agency, or the capacity to cause harm at scale. That distinction – between general capability and consequential action – is the foundation of a workable governance system.

From principles to mechanisms

AI governance scholarship has developed around four recurring concerns: fairness and discrimination, transparency and explainability, accountability, and privacy and data governance.

Research on fairness demonstrated that algorithmic systems can reproduce or amplify social inequalities even when attributes such as race or gender are not explicitly included. AI systems learn from historical data and operate inside institutions; technological neutrality cannot simply be assumed. Work on transparency asked a related question: if an algorithm influences whether a person receives a loan, a job, medical treatment, or a public benefit, how can the affected person or a regulator understand and challenge the decision?

The accountability problem is harder still. Traditional institutional arrangements generally assume that responsible human decision-makers can be identified. Automated systems complicate that assumption. A consequential decision may emerge from interactions among developers, data providers, model designers, software vendors, managers, frontline users, and automated tools. Responsibility is distributed, and can easily become diffuse.

Privacy and data governance form the fourth strand. Contemporary AI depends on large datasets, which creates questions about collection, consent, ownership, surveillance, security, secondary use, and the quality and representativeness of the information on which a model relies.

Four practical lessons follow. Governance authority is distributed among governments, firms, technical organizations, professional bodies, civil society, and international institutions. Governance must protect against harm while preserving socially valuable uses. It is inherently sociotechnical because many failures arise from the interaction of models, people, incentives, and institutional routines. And it must operate across the entire AI lifecycle, from the definition of a problem to the retirement of a system.

The unresolved question is no longer whether AI requires governance, but how to make governance work as capabilities, business models, and organizational uses continue to change. Principles are necessary, but they operate at thirty thousand feet. Leaders also need mechanisms that can diagnose, interrupt, and contain a failure before it propagates.

Why choke points are not enough

Much discussion of AI governance, especially in the United States, centers on geopolitical competition over frontier and foundation models. Policy attention has focused on advanced chips, large training runs, investment, and strategically important models.

This approach often borrows from nuclear governance. The analogy is useful, but only up to a point. Nuclear technology and AI are both transformative, can support immense benefits or severe harm, and create governance problems that cross borders. But the sources and pathways of risk are very different.

Nuclear governance was built around a comparatively tractable physical fact: the most dangerous nuclear capabilities depend on scarce materials and conspicuous industrial facilities. Uranium enrichment and plutonium production require specialized equipment, substantial investment, and supply chains that can be monitored. These characteristics created choke points at which governments could inspect, regulate, and sometimes prevent the acquisition of nuclear weapons.

A choke point suggests a narrow passage that nearly every dangerous actor must traverse. Some parts of the AI ecosystem have this character today. Production of the most advanced chips is concentrated, very large training runs are visible to a limited number of cloud providers, and frontier models are developed by relatively few organizations. These concentrations create leverage for governance. But they may be temporary, and they are poor proxies for intent.

A compute threshold cannot tell whether a model will be used to discover a medicine, automate a cyberattack, allocate public benefits, or screen job applicants. Artificial intelligence has no direct equivalent of fissile material. It’s essential inputs – data, algorithms, computing capacity, expertise, and electricity – are widely distributed and have extensive civilian uses. Models can be copied, adapted, and connected to other tools. Even when advanced models are concentrated in a few companies, their applications spread rapidly through the economy.

Upstream choke points still matter for genuinely dangerous capabilities. They can support evaluations, reporting requirements, access restrictions, export controls, and international coordination. But they cannot carry the full burden of AI governance. A model that is benign in one setting may become consequential in another because of the data, tools, permissions, authority, and scale supplied by the organization that deploys it.

From choke points to organizational control points

Instead of asking only whether AI can be controlled through nuclear-style choke points, we should also ask where, along the path from general capability to consequential action, useful control points can be established. A control point is any place where intervention can materially reduce the probability or scale of harm while preserving valuable uses. It may sit upstream, where capabilities are created; in the middle, where users gain access and systems acquire permissions; or downstream, where AI enters an organizational workflow and affects real decisions.

A simple lifecycle pathway looks like this:

Problem definition -> data acquisition -> model development -> testing -> deployment -> monitoring -> modification -> retirement.

Each stage offers opportunities for intervention. At problem definition, leaders should ask whether AI is appropriate for the task and specify what the system is meant to optimize. During data acquisition, governance should address privacy, consent, provenance, representativeness, intellectual property, and quality. During development, it should address robustness, security, explainability, foreseeable misuse, and the limits of the model’s intended use.

Testing should reproduce real operating conditions, important subgroups, adversarial circumstances, and likely workarounds rather than report only average performance. Deployment should make the division of authority between people and systems explicit. Monitoring should detect drift, unexpected use, changes in the surrounding environment, and evidence that the original assumptions no longer hold.

In the antiviral example, exploratory literature review should face little friction. Access to proprietary patient or pathogen data should face more. A model-generated recommendation to synthesize a molecule or initiate an experiment should require stronger authorization. Advancing a candidate into human trials should require stronger controls again. The model may be the same; the consequences of relying on it are not.

Most organizational failures do not begin with a single defective model. They arise in the gaps between lifecycle stages: an ambiguous objective, unrepresentative data, a test that does not match real use, an output that silently acquires decision authority, or a monitoring process that detects problems but cannot compel action. Managing those handoffs is the central task of organizational AI governance.

This lifecycle perspective is particularly important for generative and adaptive AI. Traditional software may remain relatively stable after deployment. AI systems can change as models are retrained, fine-tuned, connected to new data, given new prompts, or integrated with external tools. Governance must therefore be an ongoing process rather than a one-time approval.

Guardrails belong in the application

The phrase AI guardrail often suggests a restriction built into a model: a refusal, filter, capability limit, or access rule. Such controls matter, especially for unusually dangerous capabilities. But many of the most selective safeguards govern the context in which a system is allowed to act – what data it may use, which tools it may control, who may rely on it, what authority its output receives, and what evidence must be preserved. In these cases, the guardrail belongs chiefly to the application domain, not to AI in the abstract.

Safety-critical industries offer useful precedents. Aviation assurance ties the rigor of software development and verification to the safety consequences of an aircraft function. Drug development and pharmaceutical manufacturing already require staged evidence, documented protocols, independent review, data integrity, change control, and post-market surveillance. These regimes do not regulate software as a generic substance. They embed it within duties governing airworthiness, product quality, patient safety, and accountable operation.

AI makes those safeguards more important, not less. It can make expert assistance available to more people, compress the time between analysis and action, automate repeated decisions, and connect recommendations directly to tools. A weak control that once limited one person or one transaction can therefore fail rapidly, consistently, and at scale.

Existence on paper is not enough; enforcement is itself a control point. The Volkswagen emissions scandal is instructive because the software did not merely fail. It was designed to recognize an official test and change the vehicle’s behavior. Independent testing under real driving conditions exposed the difference between certified and actual performance. The lesson for AI is that validation cannot stop at a demonstration supplied by the developer. Application-specific regimes need independent testing, auditable records, monitoring after deployment, channels for reporting anomalies, and consequences that make institutional circumvention costly.

What AI changes in an already regulated process

Most of the control architecture needed for AI-assisted drug development already exists. Preclinical research is conducted under laboratory quality requirements. Human trials require regulatory and ethical review, defined protocols, informed consent, and safety monitoring. Manufacturing is subject to validated processes, documented changes, quality review, and batch-release requirements. Approved drugs remain subject to adverse-event reporting and post-market surveillance.

AI does not create the need for these gates. It changes the object that the gates must evaluate. A model can influence many stages at once: target selection, molecule ranking, study interpretation, regulatory documentation, trial design, manufacturing, and pharmacovigilance. A single hidden assumption may therefore appear to be independently confirmed when it is actually being repeated across the organization.

The validity of an AI model is also unusually dependent on context. A system that is useful for summarizing research may be unreliable for ranking toxicity or recommending patient doses. Governance must specify the decision the model is supporting, the population and data to which it applies, the evidence required, and the consequences of error. Model performance in the abstract is not enough.

Generative models can conceal uncertainty behind persuasive language, while predictive models can turn uncertain evidence into precise-looking scores. Reviewers therefore need access to provenance, uncertainty, missing data, subgroup performance, and plausible alternative explanations – not merely the recommendation shown on a dashboard or in a management presentation.

AI systems can also change without an obvious physical modification. A vendor may update a model; a company may change a prompt, retrieval source, dataset, or tool connection; performance may drift as the environment changes. Existing change-control principles must extend to model versions, data, prompts, integrations, and external providers.

Finally, human oversight can become nominal. A person who clicks approve is not exercising meaningful judgment if that person cannot see the model’s limitations, lacks the expertise or time to challenge it, or is rewarded for accepting its recommendation. The relevant safeguard is not the presence of a human somewhere in the workflow. It is a human with the evidence, competence, authority, and institutional support to intervene.

Graduated friction and accountability

Practical AI governance therefore requires graduated friction at control points: ordinary access for ordinary activity, with additional safeguards as a system acquires more capable tools, broader permissions, greater autonomy, or authority over consequential decisions.

The amount of friction should reflect capability, context, tools, scale, authority, and behavior – not merely the topic of a request. A diagnostic suggestion is not equivalent to authority to prescribe or deny treatment. Legal research is not equivalent to an opaque assessment that influences sentencing. A productivity assistant is not equivalent to a system that automatically screens, ranks, or terminates workers. The consequential transition is from assistance to action, and that is where stronger controls should appear.

Graduated controls can include monitoring, identity verification, bounded permissions, restricted environments, delayed execution, human authorization, independent review, impact assessment, affected-party consultation, and stronger recordkeeping. The aim is to preserve beneficial use while adding friction where information becomes organizational action or unaccountable authority.

Responsibility for control points is distributed, but it must not be allowed to become diffuse. Model developers can evaluate capabilities and manage access. Hospitals, laboratories, courts, banks, employers, and platforms understand the settings in which systems act and often already operate under domain-specific duties. Professional bodies can define competent practice. Auditors, insurers, regulators, and courts can test compliance and create consequences after failure. Leaders must connect these roles into a system in which ownership remains visible.

Controls require procedural safeguards of their own. False positives can restrict legitimate research or care. Identity requirements can disadvantage independent researchers and smaller institutions. Monitoring can threaten privacy and civil liberties, while compliance costs can entrench incumbents. Restrictions should therefore be proportionate, reviewable, and contestable. Significant decisions should be explained, subject to meaningful human review, and open to appeal.

The task for organizational leaders

For leaders, the practical task is to map where AI enters a workflow, where its outputs acquire authority, what evidence is preserved, who can interrupt the process, and what happens after a failure. The governing question is not simply, ‘Do we use AI?’ It is, ‘At which transitions can this system create harm, and who is responsible for intervening?’

This requires an inventory of uses rather than only an inventory of models. Leaders should know which systems touch sensitive data, which can call external tools, which influence consequential decisions, which can act without further approval, and which are reused across multiple stages. They should identify the assumptions that travel with a model and test whether apparently independent decisions are relying on the same underlying evidence.

They must also design escalation before a crisis. Someone needs authority to pause a deployment, study, transaction, or production process. Frontline personnel need channels for reporting anomalies without being penalized for delay. Reviewers need evidence that allows them to disagree. Monitoring needs a connection to action; a dashboard that detects a problem but cannot compel a response is observation, not governance.

Government’s role is not diminished by this approach. Sector regulators can define minimum duties, require evidence, inspect organizations, and impose consequences. Model developers can remain responsible for evaluating dangerous capabilities and providing information needed by downstream users. But the final layer of control must sit where a model is connected to a particular purpose, dataset, tool, institution, and decision.

From principles to operational control

The scientist using AI to review antiviral research and the system authorizing a laboratory experiment may rely on the same underlying model, but they do not present the same risk. Governance should preserve the first use while controlling the transition to the second.

This is neither laissez-faire nor blanket licensing of general-purpose intelligence. It is a more precise allocation of responsibility: relatively open development and ordinary access; targeted upstream safeguards for genuinely dangerous capabilities; and strong, enforceable controls when organizations give AI access, authority, physical agency, and scale.

Principles tell us what responsible AI should look like. Choke points can provide leverage over a small set of exceptional capabilities. Organizational control points make governance operational. They are where leaders can see how a model enters a workflow, where uncertainty becomes commitment, and where a preventable failure can still be stopped.

Leave a comment