1. The pilot starts with a technology, not a business decision
Many enterprise AI pilots begin with an instruction such as “find a use case for generative AI”. That reverses the order of operations. A credible pilot starts with a costly decision, delay or failure point: reducing invoice exceptions, identifying machinery faults earlier, shortening claims handling or improving demand forecasts. Without a defined operational problem, teams optimise demonstrations rather than outcomes. The result may look impressive in a workshop while remaining irrelevant to the people who control budgets and workflows.
Leaders should define the baseline, target and decision owner before selecting a model. If a contact centre currently resolves 62 per cent of enquiries on first contact, the pilot might aim for 70 per cent without increasing average handling time or complaint rates. If contract review takes four hours, the target could be 90 minutes with no decline in issue detection. This specificity forces useful trade-offs into view. Accuracy matters, but so do latency, cost, compliance and the consequences of being wrong.
2. Success is measured by model performance alone
Technical teams naturally gravitate towards precision, recall, benchmark scores and retrieval accuracy. Those measures are necessary, but they rarely determine whether a system scales. A fraud model can achieve 95 per cent recall and still fail if it creates thousands of false alerts that investigators cannot process. A document assistant may answer correctly in 85 per cent of test cases yet save no time because employees must verify every sentence against the source.
A pilot needs three layers of measurement. The first covers model quality, including accuracy, hallucination rates and performance across customer or product segments. The second measures workflow effects: minutes saved, cases completed, rework generated and adoption by intended users. The third captures financial and risk outcomes, such as margin improvement, losses avoided or regulatory exposure. A useful threshold might be a 20 per cent reduction in handling time, with critical-error rates below 1 per cent and at least 70 per cent weekly adoption after eight weeks.
Measurement should also include a control group or a credible pre-pilot baseline. Without one, teams often credit AI for improvements caused by seasonality, staffing changes or redesigned processes. A narrow test involving 100 users can provide directional evidence, but it should run long enough to reveal novelty effects. Usage frequently spikes during launch week and collapses once employees discover that the tool adds another screen, login or review step.
3. The data foundation is treated as somebody else’s problem
AI exposes data weaknesses that conventional software can often conceal. Customer records sit across incompatible systems, product descriptions are incomplete, policy documents contradict one another and access controls reflect organisational history rather than current roles. Retrieval-augmented generation cannot produce reliable answers when its source material is stale, duplicated or poorly labelled. Predictive models cannot learn a stable signal when definitions of “churn”, “profit” or “resolved case” vary between departments.
The common mistake is to treat data preparation as preliminary plumbing that should be finished quickly. In practice, it may consume 50 to 70 per cent of the effort required to make a pilot production-ready. Leaders should regard this work as part of the product: assign owners to critical datasets, document lineage, establish refresh schedules and define quality thresholds. For an employee policy assistant, that may mean identifying one authoritative version of every policy and removing expired documents before model testing begins.
There is a trade-off between waiting for perfect data and building on a fragile foundation. The answer is not a multi-year clean-up programme. Teams should identify the smallest governed data domain that can support a valuable decision, then expand deliberately. A pilot using three well-maintained product lines is more informative than one covering 40 lines with unknown gaps. Scope discipline turns data quality from an abstract transformation challenge into a measurable delivery requirement.
4. The workflow and its users are designed around the model
A model is not a product, and an interface is not a workflow. Pilots stall when teams add a chatbot beside an existing process and expect employees to reorganise their work around it. A claims handler who must copy customer details into a separate tool, inspect an answer and paste it back into the case system has not been given automation. The organisation has created an additional task, along with new opportunities for error.
Successful systems appear at the right point in the workflow and offer an appropriate level of autonomy. Low-risk actions, such as classifying inbound emails, may be automated with periodic sampling. Higher-risk actions, such as denying credit or recommending clinical treatment, require human review, explanation and escalation. The design question is therefore not simply whether a model can perform a task, but who should act on its output, what evidence they need and what happens when confidence is low.
Front-line users should shape these decisions before development, not merely test a finished prototype. Ten structured interviews and two days of observation can reveal exceptions that a process map misses: informal approval chains, peak-period workarounds and customers who require special handling. This involvement also builds trust. Training remains necessary, but it should cover changed responsibilities and error handling rather than serve as a remedy for poor design.
5. Governance arrives after the demonstration
Many pilots proceed in a low-friction sandbox, only to encounter legal, security and risk reviews shortly before launch. At that point, fundamental choices may already be fixed: customer data has been sent to an unsuitable service, prompts are not logged, vendors retain inputs, or the system cannot explain why it produced a recommendation. A promising eight-week experiment then spends six months waiting for approval or is abandoned because redesign is too expensive.
Governance should be proportional and embedded from the first sprint. The pilot owner needs to classify data, map model and vendor dependencies, define retention rules and assess foreseeable harms. High-impact use cases require testing for bias, robustness and performance drift across relevant groups. Generative systems also need protections against prompt injection, confidential-data leakage and unsupported claims. None of this requires every experiment to face the controls used for a credit decision, but the route to production must be visible.
Clear accountability is as important as technical safeguards. Someone must own the business outcome, someone must approve the risk posture and someone must operate the system after launch. A steering committee with 15 members is not a substitute for named decision rights. Organisations move faster when they publish reusable patterns: approved model providers, standard privacy clauses, logging requirements and risk tiers that determine which reviews are mandatory.
6. Teams ignore economics and production engineering
Pilot economics are often deceptive. A vendor may provide credits, engineers may absorb manual work and volumes may be too low to reveal infrastructure costs. At production scale, every document processed can incur model, retrieval, storage and monitoring charges. A system that costs 4p per interaction looks inexpensive until it handles 50 million interactions a year. Larger models may improve quality by a few percentage points while doubling latency and multiplying cost.
A realistic business case includes integration, evaluation, support, change management and human review, not just inference fees. Suppose an assistant saves five minutes on 200,000 cases annually. That represents roughly 16,700 hours, but the financial benefit depends on whether capacity is removed, demand is absorbed or service quality improves. Time saved is not automatically cash saved. Leaders should state which mechanism creates value and compare it with the full cost of ownership over two or three years.
Production engineering must also be tested early. Models change, data drifts and upstream systems fail. Teams need version control, automated evaluations, fallbacks, observability and a plan for incidents. They should test peak loads and degraded conditions rather than assume a prototype architecture will cope. A smaller model with 92 per cent task accuracy, 600-millisecond latency and predictable hosting costs may be a better enterprise choice than a frontier model delivering 95 per cent accuracy at five seconds.
7. The pilot has no owner, adoption plan or path to scale
Pilots frequently sit between innovation teams, technology functions and business units. The innovation team can demonstrate potential but lacks operational authority. Technology can build integrations but does not own the target metric. The business sponsor expresses support without assigning people or budget. When the pilot ends, nobody is accountable for converting evidence into a production decision, so the project enters an indefinite “next phase”.
Every pilot should begin with explicit stage gates. At week zero, leaders should agree what evidence will trigger termination, extension or deployment. A 12-week programme might require technical feasibility by week four, measurable workflow improvement by week eight and a signed production plan by week 12. The plan should name the product owner, operating team, funding source, integration sequence and expected roll-out cohorts. Stopping a pilot that misses its thresholds is disciplined capital allocation, not failure.
Scaling also requires organisational change. Managers may need revised targets, employees may need new quality responsibilities and control teams may need dashboards rather than periodic documents. Roll-out should proceed by repeatable units, such as one region or process type at a time, with performance compared across cohorts. The enterprises that scale AI are not those that run the most experiments. They are those that connect a sharply defined problem, governed data, usable workflow, sound economics and accountable ownership into one operating system for adoption.
Comments (0)
Discussion is opening soon. Be the first to comment.