Start with a claim finance can test
Most AI business cases begin with a percentage: 30 per cent more productive, 50 per cent faster, £2 million of annual value. The figures often come from vendor benchmarks, short pilots or employee surveys. None proves that money was saved, capacity was released or revenue increased. A credible scorecard starts by translating the claim into an observable operational change: which activity changes, for whom, at what volume, over what period and against which baseline.
Suppose a customer-service team of 120 agents introduces AI-generated response drafts. The proposal says each agent will save 45 minutes a day, implying roughly 20,000 hours a year. Finance should not price every theoretical hour at the fully loaded salary rate. It should test draft usage, average handling time, cases closed per paid hour, overtime, contractor spend, service levels and headcount plans. If handling time falls but case volumes and staffing remain unchanged, the organisation has gained capacity, not yet cash.
Every claim should therefore carry a value category. Hard savings reduce an approved budget or an external invoice. Avoided costs prevent a documented future expense. Capacity gains create usable time but require a separate conversion plan. Revenue effects need a defensible link to conversion, retention or price. Risk reduction depends on expected loss, not the dramatic cost of a single hypothetical event. This classification prevents attractive operational metrics from being presented as realised financial returns.
Measure adoption as behaviour, not access
Licences are not adoption. Neither are log-ins, training attendance or the number of employees who tried a chatbot once. The scorecard should measure repeated use within an eligible workflow: weekly active users as a share of the target population, tasks completed with AI, recommendation acceptance, edits before submission and retention after 30, 60 and 90 days. These indicators reveal whether the tool has become part of the work rather than a novelty beside it.
Adoption also needs a denominator that reflects opportunity. If 800 employees hold licences but only 300 regularly perform the supported task, reporting 240 weekly users as 30 per cent adoption is misleading; among eligible users, adoption is 80 per cent. Conversely, a high user rate can hide shallow use. Two hundred recruiters may open an assistant each week, yet use it for only 8 per cent of suitable job descriptions. Finance needs both user penetration and workflow penetration.
Segment the results by role, team and use case. A tool may deliver strong value for analysts preparing monthly reports and almost none for executives drafting occasional emails. That is not a failed programme; it is evidence for reallocating licences and implementation support. Adoption data should also expose friction. Low acceptance rates may indicate poor outputs, while heavy editing can erase time savings. Instrumentation matters because self-reported usage routinely overstates both frequency and benefit.
Build the baseline before celebrating speed
Cycle-time gains are among the most defensible AI measures, provided the baseline is fixed before deployment. Record median and 90th-percentile completion times, paid labour minutes, queue time, rework and throughput for a representative period. Use comparable work: a complex contract review should not be benchmarked against a standard renewal. Where demand or seasonality shifts, retain a control group or introduce the system in stages so that AI effects can be separated from staffing, policy and volume changes.
Consider an accounts-payable operation processing 50,000 invoices a month. An extraction model reduces median touch time from six minutes to four, apparently releasing 1,667 hours. Yet exception rates rise from 7 to 10 per cent, and each exception takes 18 minutes to resolve. The extra 1,500 exceptions consume 450 hours, reducing the net gain to 1,217 hours. If staff use 600 of those hours to clear a backlog and overtime falls by 300 hours, only the overtime reduction is immediately cashable; the remainder is capacity.
Report distributions, not just averages. AI often accelerates routine tasks while creating a long tail of difficult cases that need specialist intervention. A five-day average approval time can conceal urgent cases delayed for three weeks. Finance should see median time, tail performance, volume and unit economics together. Faster work has value only if the saved time is dependable, available where demand exists and not offset by new review or remediation effort.
Put quality beside productivity
A scorecard that rewards speed alone will encourage teams to externalise errors. Quality should be measured before and after implementation using metrics appropriate to the workflow: defect rates, first-contact resolution, factual accuracy, policy compliance, customer complaints, returns, audit findings or downstream rework. Human review is itself a cost, so the scorecard must track both the percentage of outputs reviewed and the minutes spent reviewing them.
A marketing team may increase campaign production from 40 to 65 assets a month with generative AI. If brand corrections rise from 5 to 18 per cent and legal review expands by 12 hours, gross output exaggerates the gain. Equally, a coding assistant that reduces development time by 15 per cent may still be worthwhile if escaped defects remain stable and test coverage improves. The relevant measure is quality-adjusted throughput: acceptable units completed per paid hour, not raw content, code or cases produced.
Some quality effects are positive but easy to miss. Standardised summaries can reduce variance between employees; translation tools can improve response consistency; automated checks can catch missing fields before submission. Establish thresholds rather than expecting perfection. A low-risk internal draft may tolerate minor errors with human review, while credit decisions, medical communications or regulatory filings require much stricter controls. The scorecard should make these risk appetites explicit and prevent productivity gains in one department from becoming compliance costs elsewhere.
Separate avoided cost from realised savings
Avoided cost is legitimate value, but only when tied to a credible counterfactual. If demand is forecast to rise 20 per cent and AI allows a team to absorb it without six planned hires, the value is not automatically six salaries. Finance should confirm that the roles were included in an approved plan, establish when recruitment would have occurred and subtract implementation, supervision and residual labour costs. A hiring idea that never entered the budget is not an avoided cost; it is an aspiration.
Realised savings require evidence in the ledger or workforce plan. Examples include reduced outsourcing invoices, lower overtime, cancelled software, fewer temporary staff or vacancies deliberately left unfilled. If an AI document-review system costs £180,000 a year and reduces external legal spend by £260,000, while adding £35,000 of internal review time, verified annual benefit is £45,000 before tax effects: £260,000 less £180,000 and £35,000. The larger gross saving belongs in operational analysis, not the ROI headline.
Capacity value deserves its own line. Released hours may support growth, service improvements or control work without changing expenditure. That can be strategically important, but it should be reported as hours redeployed and outcomes achieved. For example, analysts might use 3,000 released hours to review 400 additional suppliers, identifying £120,000 of duplicate payments. The verified recovery is financial impact; the hours are the enabling capacity. Keeping them separate avoids counting the same benefit twice.
Calculate verified financial impact net of full cost
AI costs extend beyond licences and model usage. Include integration, data preparation, security reviews, change management, training, monitoring, human oversight, vendor management and decommissioning. Internal labour should be valued consistently, but distinguish one-off implementation effort from recurring operating cost. If a pilot uses scarce engineering capacity, its opportunity cost should be visible even when no external invoice exists. Otherwise, apparently cheap tools are subsidised by hidden work across technology, risk and operations.
Use a simple financial bridge: gross validated benefits minus recurring costs, amortised implementation costs and quantified adverse effects equals net verified impact. Then show cash timing. A programme may generate £900,000 of annualised benefit by December but only £280,000 in the current financial year because adoption ramped slowly. Finance should resist mixing run-rate projections with booked savings. Payback period, net present value and internal rate of return can follow once the cash flows have been evidenced.
Confidence levels make the scorecard more honest. Mark ledger-backed savings as verified, benefits supported by operational evidence as probable and modelled future effects as forecast. Apply probability weightings where useful: a £500,000 retention benefit with 40 per cent confidence contributes £200,000 to a risk-adjusted view, not £500,000 to realised ROI. Keep ranges visible for volatile items such as inference costs, volumes and error rates. Precision without evidence is theatre expressed to two decimal places.
Govern the scorecard as a portfolio
The strongest governance rhythm is monthly at use-case level and quarterly at portfolio level. Each use case needs an accountable operational owner, a finance partner, a defined eligible population, a baseline, quality guardrails and a benefit-conversion plan. The owner should explain variances: adoption stalled, volumes fell, costs rose, quality breached tolerance or capacity was redirected. Finance validates classification and evidence; it should not be expected to rescue an unmeasurable project after launch.
Portfolio reporting should include adoption, workflow penetration, net cycle-time change, quality movement, released capacity, realised savings, avoided costs, revenue impact, recurring cost, cumulative cash flow and confidence status. Add stop, fix and scale decisions. A use case with 75 per cent adoption but no quality-adjusted gain may need redesign. One with 25 per cent adoption and strong unit economics may warrant targeted training. A pilot that cannot produce baseline data after two review cycles should not graduate on enthusiasm alone.
Set decision thresholds in advance. Scale when quality remains within tolerance, unit economics are positive and benefits persist for at least two reporting periods. Fix when evidence shows value but adoption or controls are weak. Stop when recurring cost exceeds plausible benefit, risk cannot be controlled or the workflow lacks enough volume. This discipline protects successful AI investments from association with inflated claims. It also gives executives a portfolio they can fund with confidence: fewer speculative totals, clearer trade-offs and financial outcomes that survive scrutiny.
Comments (0)
Discussion is opening soon. Be the first to comment.