Why AI investments need a value map
Most AI proposals begin in the wrong place. Teams compare models on benchmark scores, context windows or token prices, then leap directly to a claim about revenue or savings. The missing middle is the operating system of the business: workflows, controls, user behaviour and customer response. A model that summarises documents with 92 per cent accuracy creates no value if employees still read every document, managers distrust the output or compliance requires a full manual review.
An AI value map makes that middle explicit. It traces a chain from model capability to task performance, workflow redesign, adoption, operational metrics and, finally, financial outcomes. Each link carries assumptions that can be tested before a large contract is signed. For example, a claims insurer might connect document extraction to faster triage, a reduction in handling time from 38 to 24 minutes, 70 per cent adjuster adoption and £1.8 million in annual capacity. If any assumption fails, leaders can see exactly where the business case breaks.
This discipline also prevents false precision. A vendor may promise a 40 per cent productivity increase, but the relevant question is whether that improvement applies to an entire role or to a task occupying 15 per cent of the working week. Improving that task by 40 per cent produces a theoretical role-level gain of only 6 per cent, before accounting for review, exceptions and change management. The map turns enthusiasm into an auditable investment thesis.
Start with capabilities, not product labels
Terms such as copilot, agent and enterprise AI conceal more than they reveal. Leaders should describe the capability being purchased in operational terms: classify an inbound request, extract six fields from a contract, draft a response using approved sources, predict churn within 30 days or execute a refund under a defined threshold. Capability statements create testable boundaries. They also expose whether a general-purpose model is sufficient or whether the use case needs retrieval, deterministic rules, specialist models or human approval.
Performance must be measured against the conditions of the real workflow. A customer-service model may achieve 95 per cent intent classification on a clean benchmark but fall to 82 per cent when messages include spelling errors, mixed languages and account-specific context. That gap changes the economics. At 95 per cent accuracy, automated routing could cover most enquiries; at 82 per cent, misroutes may increase resolution time and repeat contacts. Evaluation sets should therefore include rare cases, seasonal spikes, adversarial inputs and the documents employees actually receive.
Cost and latency belong in the capability definition. A model that produces an excellent answer in 25 seconds may be unsuitable for an interactive sales call, while a cheaper model with slightly lower accuracy may be ideal for overnight invoice processing. The decision is rarely ‘best model wins’. It is a portfolio choice among quality, speed, controllability, privacy and unit cost, matched to the risk and value of each task.
Translate capability into workflow change
A capability becomes valuable only when work changes. There are four common patterns: assist, automate, augment and redesign. Assistance offers suggestions while a person retains the existing process. Automation removes a bounded task. Augmentation enables higher-quality decisions or greater personalisation. Redesign changes roles, queues and hand-offs around the new capability. Each pattern has a different investment profile. A drafting assistant can launch quickly, but a redesigned underwriting flow may unlock more value by eliminating duplicate data entry and routing only ambiguous cases to specialists.
Consider accounts payable. Extracting invoice fields is not itself an outcome. Value appears when extracted data enters the enterprise resource planning system, validates against purchase orders and sends only exceptions to a reviewer. If 60,000 annual invoices require six minutes each, full manual processing consumes 6,000 hours. An AI system that straight-through processes 65 per cent of invoices and cuts exception handling to four minutes reduces labour demand by about 4,440 hours. At a loaded cost of £32 per hour, that is roughly £142,000 in gross annual capacity before software, integration and oversight costs.
The trade-off is that deeper workflow change raises implementation risk. Integration with legacy systems, revised controls and new accountability can cost more than the model. Leaders should distinguish technical deployment from operational deployment. A chatbot accessible through a browser is technically live; it is operationally live only when triggers, ownership, escalation paths, service levels and audit records are embedded in daily work.
Measure adoption as behaviour, not access
Licences are a weak adoption metric. A company may provision 5,000 employees and report an 80 per cent activation rate even though most users tried the tool once. Useful adoption is repeated behaviour within a target workflow. Relevant signals include weekly active use, percentage of eligible tasks attempted, completion rate, acceptance or edit rate, time saved per completed task and the share of outputs requiring escalation. These measures reveal whether the tool has become part of work rather than another tab in the browser.
Adoption should also be segmented. Senior analysts may use a research assistant daily while junior staff avoid it because they cannot judge output quality. One call-centre team may achieve 75 per cent suggestion acceptance, while another reaches 30 per cent because local knowledge articles are outdated. An average obscures both the opportunity and the defect. Cohort analysis by role, manager, region and task type helps teams separate training problems from model or data problems.
Targets need a denominator tied to eligible work. If sales representatives use an email assistant 2,000 times in a month, the number sounds substantial. If they sent 40,000 eligible follow-ups, penetration is only 5 per cent. Conversely, low usage may be appropriate for a rare, high-value task such as reviewing unusual contract clauses. Adoption is not inherently good; it matters when it is associated with better cycle time, quality, risk or customer outcomes.
Connect operating metrics to financial outcomes
The financial layer should begin with a small set of operational drivers. Productivity cases typically depend on task volume, minutes saved, adoption and the percentage of released capacity that can be redeployed or removed. Revenue cases depend on reach, conversion uplift, average order value and margin. Risk cases depend on incident probability, loss severity and the degree of reduction attributable to the intervention. Writing the equation exposes assumptions that would otherwise hide inside a headline number.
Suppose an AI sales assistant serves 200 representatives, each handling 25 qualified opportunities a month. If adoption reaches 60 per cent and conversion improves from 18 to 19 per cent on assisted opportunities, the system produces about 30 additional wins monthly. At £8,000 average gross profit per win, that is £2.88 million in annual gross profit. But if the measured uplift reflects better representatives choosing to use the tool, the causal claim is unreliable. A phased rollout or matched control group is needed before the finance team books the benefit.
Costs require equal rigour. Include licences, inference, integration, data preparation, evaluation, security reviews, training, human oversight and ongoing model maintenance. Capacity savings should not automatically be labelled cash savings. Saving 10,000 hours is financially material only if the organisation avoids hiring, reduces overtime, increases throughput or reallocates staff to work with demonstrable value. Otherwise, the benefit is optional capacity, not realised profit.
Build evidence through staged investment
An effective value map supports stage gates rather than a single approval. The first stage tests capability on representative data. The second tests workflow fit with a limited group. The third measures operational impact, and the fourth validates financial value at scale. Funding should increase as uncertainty falls. A £40,000 evaluation may prevent a £1 million platform commitment built on assumptions that collapse during integration.
Each stage needs explicit thresholds. A legal review assistant might require 90 per cent recall on high-risk clauses, citations for every recommendation, median response time below ten seconds and no critical data leakage. A pilot could then require 65 per cent weekly usage, a 20 per cent reduction in first-pass review time and no increase in missed-risk incidents. If the system meets speed targets but fails recall, the team should not compensate by celebrating engagement.
Experiments must also account for novelty effects and learning curves. Early users may be unusually motivated, while productivity can initially decline as people learn new practices. Pilots should run long enough to include normal workload variation, often eight to twelve weeks, and compare outcomes with a credible baseline. Logging should capture model version, prompt or workflow version and source data so that results remain explainable after an update.
Govern the weakest links in the chain
AI risk is not confined to model accuracy. Value can fail because retrieval content is stale, permissions are too broad, staff over-trust fluent answers or suppliers change pricing and model behaviour. The value map should therefore include controls beside every link: data ownership at the capability layer, review rules in the workflow, monitoring at the adoption layer and benefit validation at the financial layer. Governance becomes part of delivery rather than a committee applied after launch.
Risk tolerance should reflect the consequence of error. A marketing tool that drafts internal headline options can accept more variability than a system recommending credit limits. High-consequence uses may require deterministic checks, dual approval, restricted actions and continuous sampling. These controls reduce theoretical automation but protect expected value. Automating 80 per cent of cases with a costly error rate may be inferior to automating 50 per cent safely and expanding as evidence improves.
Ownership must be named. Technology teams can run infrastructure and evaluations, but business leaders own workflow change and realised benefits. Finance should validate economic assumptions, risk teams should define control thresholds and frontline managers should manage adoption. A cross-functional scorecard prevents the familiar outcome in which the model performs adequately, nobody changes the process and the project is still described as a technical success.
Use the map to decide what not to fund
The most valuable result of mapping may be a rejection. Some proposals target low-volume tasks, produce benefits that cannot be captured or depend on adoption behaviour the organisation has repeatedly failed to change. Others solve a data-quality or process-standardisation problem with an expensive probabilistic layer. If a rules engine can process 95 per cent of transactions at one-tenth of the cost, generative AI should be reserved for the ambiguous remainder.
A portfolio view helps compare opportunities on expected value, evidence strength, time to impact and downside risk. A modest service-routing project with £500,000 of well-supported annual benefit may deserve priority over a speculative personalised-selling platform promising £10 million. Leaders can also identify shared enablers: improving product data may support customer service, sales and procurement use cases, producing more value than funding three isolated assistants.
Before approval, every proposal should answer five questions: what capability is required, which workflow will change, what behaviour will prove adoption, which operating metric will move and how will that movement appear in financial results? It should also state the cost, control burden and evidence needed to proceed. When those answers form a credible chain, AI investment becomes a managed business decision rather than a wager on impressive technology.
Comments (0)
Discussion is opening soon. Be the first to comment.