Scale turns AI governance into an operating problem
A pilot can survive on goodwill and informal judgement. An enterprise portfolio cannot. Once AI moves from a handful of experiments to dozens of systems embedded in sales, service, finance and operations, every output can trigger a business decision: approve a refund, prioritise a lead, flag a transaction, draft a contract or recommend a price. At that point, the central question is no longer whether the model works. It is who has the authority to define acceptable performance, who absorbs the downside when it fails and who can stop it.
Many organisations answer by creating an AI council. That may improve visibility, but a committee is not an operating model. If 60 use cases require quarterly review, the council faces 240 decisions a year before considering incidents, model changes or regulatory updates. It either becomes a bottleneck or delegates implicitly to teams without clear mandates. A scalable model must place routine decisions close to the work while reserving high-consequence decisions for accountable leaders.
The practical goal is not central control of every model. It is controlled decentralisation. Business units should own outcomes; technology teams should own shared capabilities; risk functions should set boundaries and challenge decisions; executive management should allocate capital and define appetite. Every AI system then needs a named owner, a measurable objective, an approved risk tier and an escalation route that works under pressure.
Assign one product owner to every AI-enabled decision
Each AI system should have a single product owner with authority over its lifecycle, not merely its launch. That person owns the business objective, workflow design, adoption, performance thresholds and retirement decision. For a customer-service assistant, the owner might be the head of service operations rather than the data science lead. The objective should be operational and measurable: reduce average handling time from eight minutes to six without increasing repeat contacts, complaints or compliance breaches.
Technical ownership remains distinct. Engineering teams are accountable for availability, latency, integrations, access controls and deployment quality. Model specialists monitor drift, evaluation scores and data pipelines. Vendors may provide infrastructure or foundation models. None of these parties should inherit accountability for the business decision simply because they built part of the stack. A model provider cannot decide whether a 3 per cent hallucination rate is acceptable in mortgage advice; the regulated business must make that judgement.
The product owner should maintain a decision record covering intended users, prohibited uses, fallback procedures, performance indicators and review dates. A useful rule is that one name, not a department, appears beside each critical approval. Shared ownership often means unowned tradeoffs. Contributors can be numerous, but the person authorised to accept residual operational risk must be identifiable within minutes.
Separate business accountability from independent risk challenge
The first line of defence—the business deploying AI—should own its risks. It chooses the use case, receives the benefit and controls the workflow. The second line, including legal, compliance, privacy, security and model risk, should define standards and challenge whether controls are proportionate. Internal audit, as the third line, should independently test whether the operating model works. Moving ownership to a central risk team may look prudent, but it allows business leaders to treat compliance as somebody else’s problem.
Risk classification should determine the level of oversight. A low-risk tool that summarises internal meeting notes might require standard privacy controls and sample-based quality testing. A medium-risk system that drafts customer communications could require approved templates, human review and monthly error reporting. A high-risk model influencing credit, employment, healthcare or legal rights should face independent validation, documented fairness testing, senior approval and continuous monitoring. The same generative model can sit in different tiers depending on context.
This approach also protects scarce specialists. If privacy lawyers review every prompt experiment, delivery slows without reducing material risk. If they review only after deployment, unacceptable practices become embedded. A tiered intake process should classify proposals within five working days, route standard cases through pre-approved patterns and reserve detailed scrutiny for systems with meaningful financial, safety, legal or reputational consequences. Exceptions should expire rather than becoming permanent through inertia.
Fund platforms centrally and business outcomes locally
AI funding becomes contentious when costs and benefits land in different places. A central team may pay for model access, vector databases, evaluation tooling and security controls, while a business unit captures productivity gains. Alternatively, each unit may buy its own stack, producing duplicated contracts, inconsistent controls and weak bargaining power. The answer is a split model: fund reusable foundations centrally and charge business units for use-case delivery, change management and measurable outcomes.
Consider a company with ten business units, each proposing a £300,000 assistant. Separate builds could cost £3 million before support. A shared platform costing £1.2 million, plus £120,000 per implementation, would cost £2.4 million across ten deployments. The saving is material, but only if the platform meets real needs. Central teams should publish service levels, unit costs and a roadmap; business units should be free to challenge a platform that is slower or more expensive than an approved alternative.
Investment gates should follow evidence. Discovery funding tests whether the problem is worth solving. Pilot funding establishes technical feasibility and user behaviour. Production funding requires a benefits case, control plan and accountable owner. Expansion funding depends on realised value, not demonstration quality. A project forecasting £2 million in annual savings should identify the mechanism—fewer contractor hours, lower handling time or avoided losses—and the finance function should verify whether those gains reach the profit and loss account.
Create decision rights before incidents expose the gaps
An escalation path should specify who decides, under what conditions and within what timeframe. Teams need thresholds for pausing automation, reverting to manual processing, notifying customers and informing regulators. If a pricing model produces abnormal discounts, operations may have authority to suspend it immediately, while the product owner decides whether to resume after investigation. For a high-risk system, resumption may also require approval from compliance or a designated risk executive.
Severity levels make this concrete. A level-one issue might involve a minor output error with no customer impact and a five-day remediation target. Level two could involve repeated inaccurate advice, requiring same-day containment and notification to the product owner and risk function. Level three could involve discrimination, sensitive-data exposure or material financial harm, triggering immediate suspension, executive notification and legal assessment. The thresholds should include volume and impact: 500 incorrect messages may be more serious than one conspicuous failure.
The organisation must rehearse these paths. A tabletop exercise can test a plausible scenario: a recruitment model begins systematically downgrading applicants from a particular postcode after a data refresh. Participants should determine who detects the pattern, who freezes the model, how affected decisions are identified and whether candidates are contacted. If the exercise ends with unresolved authority, the operating model has failed before the technology has.
Measure the system, not just the model
Model accuracy is only one part of performance. An AI assistant can score well in laboratory evaluations yet reduce productivity because employees spend longer checking its work. Conversely, a model with imperfect outputs may create value when it safely handles routine tasks and escalates uncertainty. Scorecards should therefore combine model quality, operational outcomes, user behaviour, control effectiveness and economics.
For a claims workflow, relevant measures could include extraction accuracy, false-denial rate, average processing time, percentage routed to human review, appeal rate, customer complaints and cost per claim. These indicators reveal tradeoffs. Lowering the confidence threshold may raise straight-through processing from 45 to 65 per cent, but if appeals double, the apparent efficiency is misleading. Metrics should be segmented by customer group, product, geography and channel so aggregate averages do not conceal concentrated harm.
Monitoring also needs decision thresholds. A dashboard without action rules merely documents deterioration. Owners should define when a metric triggers investigation, increased human review or suspension. They should also track overrides: if employees reject 40 per cent of recommendations, the model may be poorly calibrated, the interface may be weak or users may distrust automation. Each explanation requires a different intervention, and none can be solved by retraining alone.
Make the central AI function an enabler, not an empire
A central AI office adds value when it provides capabilities that business units cannot efficiently create alone: approved architectures, vendor due diligence, evaluation frameworks, model inventories, specialist advice and reusable controls. It should also convene portfolio decisions, identify duplicated investments and maintain a view of aggregate exposure. Its success is measured by safer, faster delivery across the enterprise, not by the number of projects it directly controls.
The central function should be small enough to avoid becoming a delivery queue. A hub-and-spoke structure often works: a core team owns standards and platforms, while embedded product, data and risk specialists support business domains. For every responsibility, a decision matrix should distinguish who proposes, who approves, who executes and who must be consulted. Labels such as “support” or “oversight” are too vague when a launch is delayed or an incident is unfolding.
Maturity requires periodic redesign. During the first year, central approval may be appropriate because expertise is scarce and patterns are untested. As controls become standardised, low- and medium-risk decisions can move to trained domain leaders, with audits checking adherence. High-risk deployments, novel model classes and material exceptions remain central. The strongest operating model is not the one with the most governance; it is the one that makes good decisions repeatedly, at the speed and scale the business demands.
Comments (0)
Discussion is opening soon. Be the first to comment.