Days 1–3: Define the operational problem before choosing the AI
Begin with a measurable bottleneck, not a technology shortlist. Operations and knowledge teams should identify three to five workflows where employees repeatedly retrieve information, classify requests, draft routine material or transfer data between systems. Examples include triaging 600 weekly service tickets, summarising compliance documents, preparing supplier reviews or answering internal policy questions. Record current volumes, handling times, error rates and escalation levels. A credible pilot target might be reducing average ticket triage from six minutes to two, while keeping routing accuracy above 90%.
Rank candidates using four factors: business value, feasibility, data sensitivity and frequency. High-volume, rules-rich work usually offers faster returns than rare strategic tasks. A purchasing team may benefit more from extracting fields from 2,000 monthly invoices than from asking AI to negotiate major contracts. Avoid starting with a process that is already poorly defined; automation tends to amplify ambiguity. Select one primary workflow and one fallback, then appoint an executive sponsor, a process owner, a technical lead and representatives from the employees who perform the work.
Write a one-page pilot charter covering the user group, expected outcome, exclusions and decision date. State what the system will not do. For example, it may draft a response but cannot send it, approve expenditure or make employment decisions. This boundary prevents a limited deployment from quietly becoming an ungoverned production service. It also gives procurement, security and legal teams a precise proposition to assess rather than an abstract request to ‘adopt AI’.
Days 4–7: Map the workflow and establish a baseline
Observe the work as it happens. Process documentation often omits unofficial spreadsheets, copied text, judgement calls and rework. Follow ten to 20 representative cases from arrival to completion and note every system, hand-off, approval and exception. Separate the parts that require retrieval, transformation, prediction and human judgement. A knowledge assistant, for example, may retrieve a policy, summarise the relevant clause and draft an answer, while an authorised employee remains responsible for interpretation and release.
Create a baseline using recent, representative data. Measure median and 90th-percentile handling time, because averages can conceal difficult cases. Add quality indicators such as first-time resolution, correction rates, customer complaints and hours spent searching. If a 12-person team spends 25% of its week locating information, the theoretical opportunity is 120 hours per week, but the practical saving will be lower once review, exceptions and adoption are included. Model conservative, expected and optimistic scenarios rather than promising the full theoretical gain.
Define success thresholds before the pilot begins. Useful measures include a 30% reduction in handling time, fewer than 5% materially incorrect outputs, 70% weekly active use among the pilot group and no critical security incidents. Include a stop condition: suspend the trial if the system exposes restricted information, repeatedly invents citations or causes error rates to exceed the existing process. A pilot that proves a use case unsuitable is still valuable if the evidence prevents a costly rollout.
Days 8–11: Prepare data and design the knowledge boundary
AI performance depends less on the volume of available information than on its relevance, structure and authority. Inventory the documents, records and databases needed for the chosen workflow. Label each source by owner, sensitivity, update frequency and status. Remove duplicates, obsolete procedures and drafts that conflict with approved policy. In a knowledge-base pilot, 300 current, well-tagged documents can outperform 30,000 files copied indiscriminately from a shared drive.
Decide how the system will obtain facts. Retrieval-augmented generation, which supplies selected source material to a language model at the time of a request, is usually more controllable than training a model on internal documents. It can preserve links to sources and reflect updates without retraining. However, retrieval introduces its own failure modes: poor chunking may separate a rule from its exception, weak permissions may reveal restricted text, and ambiguous searches may return a superficially relevant document. Test both straightforward and adversarial queries.
Set a minimum data standard. Documents should have clear titles, effective dates, owners and access controls; structured records should use consistent fields and definitions. Build a test set of 50 to 100 real examples, including edge cases, incomplete inputs and conflicting sources. Remove personal data where it is unnecessary, but retain enough operational complexity to make the evaluation meaningful. Synthetic examples alone tend to produce an unrealistically favourable result.
Days 12–15: Select tools against evidence, integration and control
Compare a small number of tools against the pilot charter rather than running an open-ended demonstration. Assess output quality, source citation, permission enforcement, audit logs, data residency, retention settings, model availability, administrative controls and integration effort. Ask suppliers whether customer prompts are used for model training, how subprocessors are managed and what happens to data after contract termination. Marketing claims about ‘enterprise-grade AI’ are not substitutes for contractual commitments and technical evidence.
Run the same test set through each candidate and score outputs blindly where possible. For extraction, measure field-level precision and recall. For drafting, use a rubric covering factual accuracy, completeness, tone and required review effort. Price should include licences, implementation, security assessment, employee time and ongoing monitoring. A tool costing £25 per user per month may be more expensive overall than a £60 alternative if it requires extensive manual correction or custom integration.
Choose the simplest architecture that satisfies the controls. A configured feature within an existing productivity platform may be adequate for summarisation or drafting; a dedicated system may be justified when the workflow needs complex permissions, orchestration or domain-specific evaluation. Resist premature custom development. Bespoke systems offer control but create maintenance obligations as models, APIs and regulations change. The 30-day objective is evidence of operational value, not an irreversible platform decision.
Days 16–19: Put governance and human accountability into the design
Conduct a proportionate risk review before live use. Identify personal, confidential, commercially sensitive and regulated data; map where it travels; and document who can access prompts, outputs and logs. Consider intellectual property, discrimination, records retention and sector-specific obligations. The risk profile of summarising internal meeting notes is materially different from recommending credit decisions or prioritising patients. Higher-impact uses require stronger validation, independent review and, in some cases, should not proceed as a rapid pilot.
Define human oversight as a specific operating procedure. ‘Human in the loop’ is meaningless unless the reviewer knows what to check, has time to check it and can override the system. For a policy assistant, require users to open cited sources before communicating consequential advice. For invoice extraction, automatically accept low-risk fields only above a validated confidence threshold and route exceptions to trained staff. Preserve the original input, generated output, reviewer action and final decision for auditability.
Create an incident path covering inaccurate output, sensitive-data exposure, harmful content and service failure. Name the person authorised to pause the system and specify alternative manual procedures. Update acceptable-use guidance so employees know which data may be entered, how outputs should be labelled and where concerns should be reported. Governance should enable a controlled trial, not become a stack of documents detached from daily work.
Days 20–24: Train employees through real tasks and visible limitations
Train the pilot cohort on the actual workflow rather than offering a generic seminar about artificial intelligence. A 60- to 90-minute session should cover the approved use case, data rules, prompt patterns, review requirements and escalation process. Demonstrate failures as well as successes: fabricated references, missed negations, overconfident wording and plausible but outdated answers. Employees who understand the limitations are less likely to treat fluent output as verified fact.
Provide short job aids with examples of strong inputs and acceptable outputs. Structured prompts can specify the objective, source, format and constraints: ‘Using only the attached approved procedure, list the required checks in a four-column table and cite the section for each.’ This is more reliable than a vague request to summarise a document. Yet prompt skill should not become a substitute for good product design; recurring instructions belong in templates, system configuration or workflow automation.
Recruit 10 to 30 users who represent different experience levels and attitudes, including sceptics. Give them a supported practice period and hold brief daily office hours during the first week. Monitor where they abandon the tool, rewrite outputs or revert to the old process. Adoption is not measured by log-ins alone. A system used frequently but followed by extensive correction may transfer labour rather than remove it.
Days 25–27: Deploy narrowly and measure behaviour, quality and cost
Release the tool to the defined cohort, workflow and data sources. Use staged access if the consequences of error are meaningful: start with five users, inspect the first 100 cases, then expand if thresholds are met. Keep a control sample processed through the existing method so comparisons are credible. Without a control, leaders may attribute seasonal volume changes, staffing differences or learning effects to the AI.
Capture quantitative and qualitative evidence. Track time per case, acceptance without edits, substantive correction rates, escalations, user activity, latency and cost per completed task. Review a random sample of outputs, not only those employees report as problematic. If the assistant saves four minutes on 500 weekly cases but adds one minute of review to every case, the net saving is 25 hours, not 33. At £30 per loaded staff hour, that represents roughly £39,000 annually before licence and implementation costs.
Investigate variation. New employees may gain more than experts, while complex cases may become slower because reviewers must verify subtle claims. These differences inform deployment design: the tool might be best for onboarding, first drafts or low-risk categories rather than the entire queue. Publish daily findings to the pilot team and fix obvious process defects quickly, but avoid changing models, prompts and evaluation criteria simultaneously; otherwise, the evidence becomes difficult to interpret.
Days 28–30: Decide whether to scale, revise or stop
Compare results with the thresholds in the pilot charter. Present benefits alongside errors, operational dependencies and unresolved risks. A useful decision paper should include baseline and pilot performance, annualised economics, employee feedback, security findings and the work required for production. Distinguish cashable savings from capacity released. Saving 20 hours per week does not automatically reduce expenditure, but it may allow the team to absorb growth, shorten response times or redirect effort to higher-value analysis.
Choose one of three outcomes. Scale when quality, controls and economics are convincing; revise when the use case is sound but data, workflow or training needs improvement; stop when risk or review effort outweighs value. Scaling should remain phased. Expand to one adjacent team or document set, establish service ownership, schedule model and retrieval evaluations, and define quarterly reviews for accuracy, access and cost. Supplier model changes can alter performance even when the interface appears unchanged.
The first 30 days should produce an operating capability, not merely a demonstration. Leaders should leave with a tested workflow, governed data sources, trained users, measurable results and a decision record. That foundation makes subsequent adoption faster and safer: each new use case can reuse the evaluation method, risk controls, training assets and procurement standards while still being judged on its own operational merits.
Comments (0)
Discussion is opening soon. Be the first to comment.