Why AI portfolios need a harder audit
Most organisations now have enough AI activity to create a portfolio problem. A retailer may be testing product descriptions, demand forecasting, customer-service copilots and computer vision at the same time. A bank may have dozens of pilots spread across fraud, compliance, software engineering and relationship management. Yet funding decisions are still too often driven by executive enthusiasm, vendor demonstrations or the visibility of a particular team. The result is predictable: promising systems remain trapped in pilots, weak initiatives absorb scarce engineering capacity, and apparently cheap generative AI tools accumulate substantial operating costs.
An AI portfolio audit replaces anecdote with comparison. Its purpose is not to rank every initiative by model accuracy or novelty, but to decide which use cases deserve capital, remediation or closure. Each initiative should be assessed on four dimensions: business impact, technical feasibility, adoption risk and total operating cost. A fifth question sits above the scorecard: does the use case remain strategically relevant? A perfectly executed tool that automates a process the company intends to retire is still a poor investment. The audit should therefore combine quantitative evidence with a clear view of operating strategy.
Build a scoring model that exposes trade-offs
A practical model scores each dimension from one to five, then applies weights that reflect the organisation’s priorities. A balanced starting point is 35 per cent for business impact, 25 per cent for technical feasibility, 20 per cent for adoption risk and 20 per cent for total operating cost. For adoption risk and cost, five means low risk or attractive economics. A use case scoring five, four, three and three respectively receives 3.95 out of five. That is useful, but the composite number should never conceal a fatal weakness. Regulatory non-compliance, unsafe outputs or unavailable data should operate as gates, not merely reduce the average.
Scoring criteria must be observable. Business impact can include annual cash benefit, revenue protection, cycle-time reduction and strategic reach. Technical feasibility covers data quality, integration complexity, model performance and production resilience. Adoption risk includes workflow fit, user trust, incentives, training and the consequences of an error. Total operating cost should include inference, licences, data pipelines, monitoring, human review, security, support and model upgrades. Define what each score means before reviewing initiatives. If a business-impact score of five requires more than £5 million in annual value, do not award it to a pilot supported only by positive user comments.
The same weights need not apply everywhere. A regulated insurer may raise adoption and control risk to 30 per cent. A digital marketplace facing a narrow competitive window may give impact and speed greater weight. What matters is consistency within the decision set. Comparing a fraud model, an employee chatbot and a warehouse-vision system is already difficult; allowing sponsors to invent their own scoring logic makes the exercise meaningless. A central review team should calibrate scores using two or three reference cases before assessing the full portfolio.
Measure business impact against a credible baseline
AI benefits should be calculated against the next-best alternative, not against doing nothing forever. Consider a customer-service copilot that reduces average handling time from eight minutes to seven. Across 500 agents handling 35 calls a day for 220 working days, the gross capacity gain is about 64,000 hours. At a fully loaded labour cost of £28 an hour, that appears to be worth £1.8 million. But if demand is flat and staffing cannot be reduced or redeployed, much of that benefit is theoretical. A credible case states whether the gain becomes lower overtime, higher service levels, avoided recruitment or additional sales.
Revenue claims require similar discipline. A recommendation engine that lifts conversion by 2 per cent in an experiment has not necessarily created 2 per cent more company revenue. The audit must account for substitution, discounting, returns and customers who would have purchased anyway. Controlled trials, phased roll-outs and matched cohorts provide stronger evidence than pre-and-post comparisons. Where measurement remains immature, apply a confidence discount. A projected £4 million benefit with 50 per cent evidence confidence should not outrank a validated £2.5 million saving simply because its headline is larger.
Impact also includes risk reduction, but risk should be monetised carefully. A compliance model that detects 20 per cent more suspicious transactions may prevent losses or regulatory penalties, even if the exact benefit is uncertain. Use expected value: probability of an event multiplied by its financial consequence, adjusted for the model’s incremental effect. Then record non-financial benefits separately, such as faster investigations or improved auditability. This prevents vague claims about ‘better decisions’ from receiving the same weight as measurable cash outcomes.
Test technical feasibility in production, not the laboratory
Prototype performance is a weak proxy for production feasibility. A document-extraction model may achieve 94 per cent field accuracy on a curated test set, yet fail when scans are rotated, forms change or handwriting appears. The relevant question is whether the entire system can deliver the required service level under real operating conditions. That includes data availability, latency, integration, security, observability, fallback procedures and the rate at which inputs drift. For high-volume processing, even a 2 per cent exception rate can create an expensive manual queue.
Generative AI demands task-specific evaluation. General benchmarks say little about whether a model can draft a compliant mortgage letter or diagnose an equipment fault from service notes. Build a representative evaluation set, define unacceptable errors and measure both quality and consistency. A legal summarisation tool might score well overall but still be unfit if it omits a liability clause in one document out of 200. Teams should also test smaller or specialised models. If a model costing £0.01 per transaction achieves 92 per cent acceptable outputs and a premium model costing £0.08 reaches 94 per cent, the extra two points may not justify an eightfold inference bill.
Feasibility scores should reflect remediation effort, not optimism. Missing application programming interfaces, fragmented ownership and poor data lineage are engineering liabilities with budgets and timelines. Estimate the work required to reach production readiness and identify dependencies shared across use cases. Sometimes the right portfolio decision is to fix a platform constraint rather than cancel several initiatives. A governed retrieval layer, for example, may unlock multiple knowledge assistants more economically than each business unit building its own document pipeline.
Treat adoption as an operating-model challenge
Many AI initiatives fail after deployment because they add friction or threaten established roles. A sales copilot that produces useful account briefs can still be ignored if representatives must leave the customer relationship system, wait 40 seconds and verify every claim. By contrast, a modest model embedded directly into the opportunity screen may gain rapid adoption. The audit should measure weekly active use, completion rates, override rates, time saved and the proportion of outputs requiring correction. Survey enthusiasm is secondary to observed behaviour.
Adoption risk rises with the consequence of error and the ambiguity of accountability. An internal search assistant can tolerate occasional irrelevant answers; a system recommending cancer treatment cannot. Human review is not automatically a solution. If reviewers approve 98 per cent of outputs, attention decays and oversight becomes ceremonial. The portfolio review should ask who owns each decision, what evidence the user sees, how challenges are recorded and whether workloads permit meaningful review. It should also identify incentive conflicts: a claims handler measured on speed may accept an AI recommendation too readily, while one punished for errors may avoid the tool altogether.
The strongest adoption plans redesign work rather than bolt a model onto it. One European manufacturer cut maintenance-planning time by 30 per cent only after standardising fault codes and changing supervisor routines; the model itself had been available months earlier. Budget for process design, training, communications and local support. If these costs make the economics unattractive, that is not a change-management failure to hide. It is evidence that the use case should be fixed, narrowed or stopped.
Calculate the full cost of keeping AI alive
AI business cases routinely understate cost by focusing on pilot development or model access. Total operating cost includes cloud infrastructure, token usage, vector databases, data acquisition, integration maintenance, security reviews, evaluation, monitoring, human escalation, vendor management and periodic retraining. A chatbot serving 100,000 conversations a month might incur only £12,000 in model charges, yet cost £250,000 a year once support, content maintenance, quality assurance and compliance oversight are included. The denominator matters too: cost per resolved enquiry is more informative than cost per generated response.
Costs also change with scale. Some decline as fixed platform investments are shared; others rise sharply because more users produce more exceptions, security exposure and demand for support. Run scenarios for expected, high and low volumes, and include sensitivity to model prices and usage patterns. A long conversational context can multiply inference cost without improving outcomes. Caching, routing simple tasks to smaller models and limiting retrieval to relevant documents can materially improve unit economics. These are architectural choices, not procurement footnotes.
Vendor concentration and exit costs belong in the calculation. A low introductory licence can become expensive when an application depends on proprietary workflows, embeddings or fine-tuned models. Estimate the cost of migration and require access to logs, evaluation data and configuration artefacts. For critical use cases, resilience may justify paying for a second provider or an internal fallback. The cheapest system in steady state is not necessarily the lowest-cost portfolio choice once switching risk and service interruption are priced in.
Decide what to scale, fix or shut down
Create decision bands, but combine them with explicit gates. Initiatives scoring 4.0 or above, with validated benefits and no red flags, are candidates to scale. Those between 3.0 and 3.9 should receive a time-bound remediation plan. Anything below 3.0 should normally stop unless it provides a strategic capability that cannot yet be valued conventionally. A high score does not authorise immediate enterprise deployment: scale in stages, with funding released when adoption, quality and unit-cost thresholds are met. A claims assistant, for example, might expand from 50 to 300 users only after maintaining 95 per cent citation accuracy and saving ten minutes per case for eight weeks.
The fix category needs discipline because it can become a graveyard of permanent pilots. Specify the defect, owner, budget and deadline. A forecasting model with strong potential but poor adoption might receive 90 days to integrate into the planning system and reach 70 per cent active use. A knowledge assistant with high hallucination rates might be narrowed to five approved document collections and retested against 500 questions. If the initiative misses its remediation threshold, close it. Sunk development cost is not a reason to continue spending.
Shutdowns should preserve learning. Archive evaluations, prompt designs, data contracts, user research and post-mortems so that another team does not repeat the same experiment. Remove dormant integrations, revoke access and terminate licences; abandoned AI services create security and cost exposure. Communicate that closure is portfolio management, not punishment. Leaders who celebrate launches but stigmatise shutdowns encourage teams to manipulate evidence and keep weak systems alive.
Make the audit a recurring capital-allocation process
A portfolio audit should run at least quarterly for fast-moving generative AI and twice yearly for more stable analytical systems. The review group needs business, finance, technology, data, security, legal and frontline representation. Product owners submit standard evidence: current usage, realised value, quality metrics, incidents, operating cost, dependencies and the next funding request. Finance should validate benefits, while risk functions test controls without assuming ownership of the business decision. A one-page scorecard for each use case keeps debate focused on evidence.
Track portfolio-level concentration as well as individual performance. Ten assistants built on the same model provider may look healthy separately while creating a serious dependency collectively. Multiple teams may also be paying to solve the same retrieval, identity or monitoring problem. The audit should identify reusable capabilities and retire duplication. A company spending £6 million across 40 pilots may discover that five production platforms, funded properly, offer more value than another cycle of disconnected experimentation.
The final output is a reallocation plan: capital and engineering capacity move towards proven use cases, constrained initiatives receive targeted fixes, and low-value work ends. Publish decisions, assumptions and review dates. Scores will change as model prices fall, regulations evolve and users gain experience; that is a feature, not a flaw. The objective is not to produce a permanent ranking. It is to make AI investment responsive to evidence, exposing where value is real, where execution can improve and where organisational attention is better spent elsewhere.
Comments (0)
Discussion is opening soon. Be the first to comment.