Start With the Decision, Not the Demonstration
AI procurement often goes wrong before a vendor enters the room. A team buys an impressive capability, then searches for a business problem that justifies it. Reverse that sequence. Define the decision or workflow being improved, the current baseline and the acceptable failure rate. For a customer-service assistant, record handling time, first-contact resolution, escalation rate and quality scores before testing automation. If agents currently resolve 78 per cent of cases and the proposed system reaches 82 per cent only by doubling escalations, the headline improvement is misleading.
Ask every vendor to describe the exact operating conditions under which its claims were measured. Was a reported 30 per cent productivity gain observed in production, a controlled pilot or an internal demonstration? How many users, documents and weeks were included? Which human checks remained in place? Demand results for customers with comparable languages, regulation, data quality and transaction volumes. A model tested on clean English-language policies may deteriorate sharply when faced with scanned invoices, regional terminology or multilingual queries.
Set acceptance criteria before the proof of concept. These should include quality, latency, cost and operational limits: for example, at least 90 per cent extraction accuracy on a representative 5,000-document sample, responses within three seconds at the 95th percentile, and a maximum cost of 8 pence per completed transaction. Pre-agreed thresholds prevent a polished demo or selectively chosen examples from redefining success after the fact.
Interrogate Model Quality and Failure Modes
“Which model do you use?” is necessary but insufficient. Ask whether the product relies on a proprietary model, an open-weight model, a routed mixture or several models assigned to different tasks. Establish model and version names, release dates, context limits, supported languages, fine-tuning methods and the vendor’s process for replacing them. If an upstream provider changes pricing or retires a model, will performance, cost or availability change without your approval? A stable product needs more than a fashionable model name.
Request evaluation results on your data and insist on task-specific measures. Accuracy may suit classification, but retrieval systems also need precision and recall, while generative applications require groundedness, citation accuracy and human assessment. A vendor claiming 95 per cent accuracy should disclose the sample size, class balance and confidence interval. On a dataset where 95 per cent of items are routine, a system can reach 95 per cent by failing every exceptional case—the cases most likely to carry financial or safety risk.
Failure analysis matters more than the average score. Ask for the ten most common error categories, examples of high-confidence wrong answers and controls for hallucination, prompt injection and out-of-scope requests. Determine when the system abstains, how uncertainty is communicated and whether users can verify sources. In a legal drafting tool, an unsupported quotation is not merely a poor response; it can contaminate advice. Buyers should test adversarial prompts, ambiguous instructions, stale documents and deliberately conflicting sources, not just happy paths.
Trace Every Piece of Data
Vendors should provide a data-flow diagram showing what enters the service, where it is processed, where it is stored, which subprocessors receive it and when it is deleted. Distinguish prompts, uploaded files, retrieved records, outputs, telemetry, feedback and support logs. “We do not train on customer data” does not answer whether data is retained for 30 days, reviewed by contractors or transferred outside the UK or European Economic Area.
Ask whether customer content is used for model training, product improvement, abuse monitoring or evaluation, and whether each purpose is opt-in or opt-out. Contractual language should override website policies that can change unilaterally. Establish data-residency options, encryption standards, backup locations and deletion timelines, including derived artefacts such as embeddings, caches and fine-tuned weights. Deleting a source document while leaving its vector representation accessible is not complete deletion.
The vendor must explain its legal basis for processing personal data and its support for data-subject rights, retention schedules and impact assessments. If special-category or children’s data could appear, require explicit controls rather than assurances that users should not submit it. Also investigate provenance: what data trained the underlying model, how licensing and copyright risks are addressed, and whether outputs can reproduce protected material. Perfect transparency may be impossible for frontier models, but uncertainty itself is a procurement risk that should be priced and allocated.
Treat Security as an Architecture Question
A security certificate is evidence, not a substitute for scrutiny. Review ISO 27001 or SOC 2 scope, report dates, exclusions and significant findings. Confirm single sign-on, multifactor authentication, role-based access, audit logs, tenant isolation, key management and secure development practices. Ask whether administrators can restrict connectors, exports and model choices. A tool with strong infrastructure security can still leak confidential information if any employee can connect it to an unrestricted public workspace.
AI introduces attack paths that conventional questionnaires often miss. Require controls for prompt injection, poisoned retrieval content, malicious file uploads, model extraction and sensitive-data disclosure. If an assistant can browse websites or execute actions, determine how domains, permissions and transaction values are constrained. A procurement bot that may draft an order is materially safer than one authorised to submit a £100,000 purchase without independent approval.
Demand an incident-response commitment covering notification times, investigation, evidence preservation and remediation. “Without undue delay” may be legally familiar but operationally weak; a contractual target such as 24 hours after confirming an incident creates clearer accountability. Review penetration-test summaries and vulnerability-disclosure processes, then identify the weakest dependency. A vendor may run a mature platform while relying on a small model gateway, transcription service or browser extension with lower standards.
Test Portability Before Lock-In Begins
AI products accumulate dependency quickly. Prompts, evaluations, embeddings, feedback, workflow logic and fine-tuned models can become more valuable than the original software subscription. Ask which assets belong to the customer, which can be exported and in what formats. An export containing PDFs but not metadata, annotations or evaluation histories may preserve documents while destroying the operating system built around them.
Test whether the product can switch models without rebuilding integrations. Model abstraction offers flexibility, but it can also hide differences in safety controls, context handling and output quality. Require notice before material model changes and the right to retest critical workflows. For high-risk uses, seek version pinning or a defined migration window. If a vendor replaces a model on Friday and accuracy falls five percentage points on Monday, operational teams need both evidence and recourse.
Exit planning should cover APIs, rate limits, export assistance, deletion certification and service continuity. Ask how long data remains accessible after termination and what professional services will cost. Consider escrow or continuity arrangements where the system supports essential operations. Portability is not necessarily free: a managed proprietary model may outperform a portable open-weight alternative. The buyer’s task is to understand the premium and avoid discovering switching costs during a dispute.
Model the Full Economics, Not the Licence
AI pricing can combine user licences, tokens, documents, API calls, storage, connectors, fine-tuning, premium models and support. Build scenarios for normal, peak and adverse usage. A £30-per-user monthly licence appears predictable until a retrieval feature adds metered embedding and generation charges. Conversely, token pricing may look volatile but cost less for occasional users than buying 2,000 seats that remain idle.
Ask the vendor to price a representative workload and state every assumption: input and output tokens, retries, context size, cache hits, document pages and currency conversion. Include implementation, data preparation, integration, security review, training, human oversight and ongoing evaluation. If a system saves five minutes per case across 100,000 annual cases, it releases roughly 8,333 hours—but only if staff can use that capacity and error correction does not consume the gain.
Negotiate budget controls, usage dashboards, alerts, volume tiers and protection against abrupt price changes. Clarify whether failed requests, automated evaluations and vendor-initiated model migrations are billable. Seek unit economics tied to an outcome, such as cost per correctly processed invoice, rather than cost per token. A cheaper model with an 8 per cent rework rate can be more expensive than a premium model that reduces manual review to 2 per cent.
Verify Support, Governance and Operational Readiness
Support promises must match business impact. Identify service hours, response and restoration targets, escalation routes and named responsibilities. A four-hour response is inadequate if the service blocks a daily payment run; it may be generous for an optional writing assistant. Ask for historical uptime, severity definitions and exclusions, not merely a 99.9 per cent headline. That figure still permits about 43 minutes of downtime per month, and maintenance windows may sit outside it.
Examine how the vendor monitors model drift, harmful outputs and changes in user behaviour. Buyers should receive release notes, performance reports and controls to pause features. Define who approves new use cases, reviews high-risk decisions and investigates complaints. Logs must support reconstruction: which model, prompt, sources, settings and user actions produced an output? Without traceability, neither the vendor nor the customer can diagnose a disputed decision.
Reference checks should include operational counterparts, not only enthusiastic executives supplied by sales. Ask customers what broke after launch, how long integration took, whether costs matched forecasts and how the vendor responded under pressure. Also review the supplier’s financial position, staffing, dependency on upstream providers and product roadmap. A technically strong service can still become an operational liability if its support team is thin or its funding horizon is short.
Put Risk Allocation Into the Contract
The contract should convert due-diligence answers into enforceable obligations. Attach security measures, data locations, subprocessors, service levels, model-change procedures and acceptance criteria. Ensure the data-processing agreement reflects actual architecture. Sales presentations and questionnaire responses should be incorporated where material; otherwise, assurances about non-training, deletion or residency may disappear behind a broad entire-agreement clause.
Allocate intellectual-property and regulatory risk deliberately. Clarify ownership of inputs, outputs, prompts, configurations, fine-tuned assets and feedback. Seek warranties that the vendor has rights to provide the service, alongside indemnities appropriate to infringement, confidentiality breaches and unlawful processing. Vendors may resist unlimited exposure, particularly for customer-directed prompts. A balanced position distinguishes risks controlled by the supplier from misuse, prohibited data and unreviewed deployment controlled by the buyer.
Liability caps, termination rights and audit provisions should reflect plausible harm rather than annual fees alone. A £50,000 cap may be commercially neat but inadequate for a system handling millions of customer records. Require cooperation with regulatory enquiries, deletion or return of data at exit, and remedies for repeated performance failures. Finally, maintain an internal decision record documenting evidence, residual risks, accountable owners and review dates. Procurement is not completed at signature; model updates, new connectors and expanded use can change the risk profile within weeks.
Comments (0)
Discussion is opening soon. Be the first to comment.