The contract, not the demo, defines the risk
Enterprise AI purchases are often approved on the strength of a polished demonstration, a security questionnaire and a standard software-as-a-service agreement. That process misses the point. An AI service does more than store and process information: it may retain prompts, infer sensitive attributes, route data through several model providers, change behaviour without notice and generate content that creates regulatory or intellectual-property exposure. If the contract treats the product like conventional workflow software, the customer inherits risks that were never priced into the deal.
Legal and procurement teams should therefore map each material promise in the sales process to an enforceable clause. A vendor may say customer data is ‘not used for training’, but the contract might permit use for service improvement, abuse monitoring or development of aggregated insights. It may promise 99.9 per cent availability while excluding the underlying model provider’s downtime. It may describe explainability and audit tools that are absent from the service schedule. Marketing copy, trust-centre statements and verbal assurances can change; negotiated contractual obligations cannot be replaced as easily.
Data retention needs dates, locations and deletion evidence
A clause stating that the supplier retains data ‘only as necessary’ is not a retention policy. The agreement should distinguish prompts, uploaded files, outputs, embeddings, conversation histories, system logs, human-review samples and backups. Each category needs a defined retention period and purpose. A practical schedule might require production content to be deleted within 30 days, security logs within 180 days and inaccessible backups within 90 days after termination. The customer should also know where each copy is stored and whether support personnel can access it from another jurisdiction.
Deletion must cover the entire delivery chain. If a vendor sends prompts to a foundation-model provider, uses a separate vector database and relies on a logging platform, deleting the visible account may leave several residual copies. Contracts should require the supplier to flow deletion obligations to subprocessors, prevent restoration except for disaster recovery, and provide written certification after deletion. Legal holds and regulatory retention can be valid exceptions, but they should be narrowly defined, disclosed and subject to access restrictions.
The trade-off is operational. Very short log retention can make fraud investigations and incident diagnosis harder, while immediate backup deletion may be technically unrealistic. Buyers should not demand impossible guarantees; they should demand specificity. A documented 90-day backup cycle is more useful than an absolute promise contradicted by the supplier’s architecture. Where data is especially sensitive, such as patient notes or unreleased financial results, the contract should support zero-retention processing or a segregated deployment at an agreed price.
Training restrictions must close the ‘service improvement’ loophole
The most important data-use clause is often buried in the licence rather than the privacy schedule. It should prohibit the supplier and its subprocessors from using customer inputs, outputs, metadata or feedback to train, fine-tune, evaluate or benchmark models unless the customer gives explicit, case-specific consent. ‘Training’ alone is too narrow: vendors may use data to build classifiers, create synthetic datasets, test new releases or employ human reviewers under the label of quality assurance.
An enterprise may reasonably allow limited use of telemetry to improve reliability, but the permitted fields and purposes should be listed. Token counts, latency and error codes pose different risks from full prompt text. Any claim that information is anonymised should be backed by a defined standard, controls against re-identification and a ban on combining it with other datasets. Aggregation is not a magic word; ten unusual prompts from one pharmaceutical research team can remain commercially revealing even after names are removed.
Feedback deserves separate treatment. Staff frequently click thumbs-up buttons, correct generated text or submit examples to support. Those actions should not silently grant the vendor a perpetual licence to valuable know-how. The agreement can permit use of voluntarily submitted feedback while excluding customer content, confidential information, personal data and intellectual property embedded within it. For regulated deployments, the safest default is no model improvement use without a written amendment.
Audit rights must reach the AI supply chain
A current ISO certificate or SOC 2 report is useful, but it rarely answers AI-specific questions: which model processed a prompt, whether content filters were active, how evaluations were conducted or whether the vendor changed its retrieval pipeline. Audit clauses should provide access to relevant policies, architecture summaries, penetration-test results, model and subprocessor registers, evaluation reports and incident records. Customers in high-risk sectors may also need evidence supporting impact assessments and regulatory enquiries.
The right must be proportionate enough to survive negotiation. Suppliers will resist unlimited on-site inspections because they create security and confidentiality risks. A sensible structure begins with independent assurance reports and written responses, escalates to a remote audit when material gaps remain, and permits an on-site or third-party audit following a serious incident or regulator request. Costs can sit with the customer for routine audits and shift to the supplier if a material breach is found.
Subprocessors are the weak link. An AI application vendor may depend on two model providers, a cloud platform, a content-moderation service and data-labelling contractors. The contract should require an up-to-date list, advance notice of material additions and a genuine objection mechanism. If the supplier cannot remove a disputed subprocessor, the customer should be able to terminate the affected service without penalty. Audit evidence should cover these providers rather than ending at the vendor’s corporate boundary.
Availability promises should cover degraded intelligence, not just uptime
Traditional service-level agreements measure whether an endpoint responds. An AI endpoint can return HTTP 200 while producing unusable answers, timing out after 50 seconds or operating without retrieval and safety controls. The agreement should define availability across critical functions, including authentication, generation, retrieval, moderation and administrative controls. It should also set latency targets by use case and identify rate limits that could restrict production volumes during peak periods.
Dependencies must not disappear into exclusions. If a supplier chooses a particular foundation model, an outage at that provider is part of the service risk it is selling. Blanket exclusions for third-party failures can make a 99.9 per cent commitment meaningless. Buyers should seek service credits for missed targets, chronic-failure termination rights and a continuity plan that addresses model-provider outages. Three breaches in a rolling six-month period, for example, might trigger termination without early-exit charges.
Fallback models create their own trade-offs. Routing traffic to a smaller model can preserve availability but reduce accuracy, change data residency or introduce a different licence. The contract should specify approved fallback providers, minimum security controls and whether the customer can disable automatic failover. Material degradation should be disclosed through status notifications and logs. For safety-critical workflows, graceful failure may be preferable to a lower-quality answer presented with the same confidence.
Indemnities must reflect how AI claims actually arise
Standard supplier indemnities often cover third-party claims that the software itself infringes intellectual property. That may not protect a customer when generated text, code or images allegedly copy protected material. The clause should expressly address outputs created through authorised use, including defence costs, damages and settlements. It should also explain exclusions for customer-supplied material, prohibited prompts, post-generation modifications and use after an infringement warning.
The allocation should track control. A supplier that selects training data, model architecture and output filters is better placed to manage model-level copyright risk. A customer that asks the system to imitate a living artist or reproduce a competitor’s source code should bear more responsibility. Neither side should accept unlimited exposure for conduct it cannot control. Caps can be higher for confidentiality breaches, data-protection failures and IP claims than for ordinary service failures; two to five times annual fees is a common negotiating range, although the appropriate figure depends on potential loss.
Indemnity is not a substitute for compliance. Contracts should require the supplier to maintain suitable insurance, notify claims promptly, preserve evidence and avoid settlements that admit customer wrongdoing or restrict operations without consent. Buyers should also examine defence conditions: an indemnity that applies only after a final judgment may provide little practical value because most disputes settle. Separate treatment may be needed for biometric, discrimination or consumer-protection claims, particularly where the system influences employment, credit or healthcare decisions.
Portability must include context, configuration and transition support
Exporting prompts and outputs as a CSV file is not meaningful portability. An operational AI system may depend on embeddings, vector indexes, retrieval documents, system prompts, agent instructions, evaluation datasets, safety rules, fine-tuning artefacts, user permissions and audit logs. The exit schedule should list exportable assets, formats, delivery times and charges. Open formats such as JSON, CSV and common document types are preferable to opaque proprietary packages.
Some components cannot be transferred directly. Model weights may be shared across customers, and embeddings generated by one model may not perform correctly with another. The contract should identify these constraints before purchase and require reasonable transition assistance, such as document re-export, configuration documentation and continued read-only access for 60 or 90 days. Fees should be pre-agreed or tied to published professional-services rates, preventing exit costs from becoming leverage.
Portability also needs an operational trigger. Customers should be able to extract data during the term, not only after termination, and the supplier should not withhold exports during a good-faith payment dispute. A tested annual export is more valuable than a theoretical right. Procurement teams can make that test an acceptance criterion for strategically important systems, particularly where AI is embedded in customer support, underwriting or internal knowledge management.
Silent model changes require notice, testing and a right to refuse
AI suppliers update models frequently to improve capability, reduce cost or respond to safety concerns. A new version may nevertheless alter tone, accuracy, refusal rates, latency, token limits or regional processing. For a bank using generated summaries in compliance reviews, even a modest shift can invalidate validation evidence. Contracts should distinguish routine patches from material changes and require advance notice, release notes and updated documentation for the latter.
Customers need a controlled adoption path. The service should provide model version identifiers in logs, a test environment and, where feasible, a 30- to 90-day period before mandatory migration. The buyer should be able to run its own evaluation set against the proposed version and reject a change that materially reduces agreed performance or compliance. Emergency security updates may require shorter notice, but the supplier should explain the reason and provide remediation options.
Performance commitments must be measurable without pretending that generative systems are deterministic. Instead of guaranteeing that every answer is accurate, parties can agree task-specific thresholds: extraction precision on a defined test set, maximum unsupported-citation rates, language coverage or human-escalation performance. If a change breaches those thresholds, the supplier should restore the previous version, remediate within a fixed period or permit termination of the affected service. Without those rights, the enterprise has not bought a stable product; it has rented access to an experiment whose rules can change overnight.
Comments (0)
Discussion is opening soon. Be the first to comment.