Prompting is not the bottleneck
Companies often diagnose disappointing AI adoption as a skills problem: employees need better prompts, more training or a library of approved instructions. That diagnosis is attractive because it turns a difficult organisational problem into a manageable learning programme. Yet most failures occur before anyone types into a model. Teams have not specified which decision the system should support, what evidence matters, who owns the outcome or how errors will be detected. A polished prompt cannot compensate for an undefined job.
Consider a marketing team asked to ‘use AI for campaign content’. The instruction leaves unresolved whether the objective is faster drafting, more variants, stronger conversion or lower agency spend. It also says nothing about audience evidence, brand constraints, legal approval or performance feedback. Staff may produce competent copy and still create no measurable value. The real skill is workflow design: breaking work into decisions, assigning responsibility and connecting outputs to consequences. Prompt craft matters, but it is one component in that larger operating system.
Start with the decision, not the model
A useful AI workflow begins with a specific decision. A sales team does not need a generic assistant that ‘analyses accounts’; it needs to decide which 50 of 2,000 prospects should receive attention this week. That framing exposes the inputs required: recent product usage, buying signals, account size, previous contact and territory rules. It also establishes an output that can be evaluated against pipeline created, meetings booked or conversion to opportunity.
Decision-first design prevents a common form of automation theatre. Summaries, chat interfaces and generated reports look productive, but they often add another layer for employees to read without changing what happens next. If an account brief does not alter prioritisation, messaging or timing, it is information without operational leverage. Each AI output should therefore have a named consumer, a deadline and a downstream action.
This approach also clarifies where automation should stop. A model might rank prospects and explain the signals behind each recommendation, while an account executive makes the final choice because they know about an unrecorded procurement freeze. Full automation could save more time, but assisted judgement may produce better commercial outcomes. The right boundary depends on the cost of errors, the quality of available data and the reversibility of the decision.
Decompose work into verifiable steps
Large, vague assignments encourage models to improvise. Asking AI to ‘write a market-entry strategy’ combines research, source selection, competitor analysis, customer segmentation, financial assumptions and executive communication in one opaque request. The result can sound authoritative while concealing weak evidence. A stronger workflow decomposes the task: collect approved sources, extract claims, compare competitors, identify gaps, test assumptions and only then draft the recommendation.
Decomposition creates checkpoints. A human can reject unreliable sources before they contaminate the analysis, or challenge a revenue assumption before it appears in a board paper. It also makes failure diagnosable. If the final recommendation is poor, the team can determine whether retrieval missed relevant documents, extraction distorted facts, criteria were incomplete or synthesis overreached. Without such visibility, every problem is blamed on the model or the prompt.
There is a tradeoff. More stages add latency and maintenance, and not every low-risk task warrants an elaborate pipeline. A two-minute internal email may need only a draft and a quick review. A customer credit decision, regulatory submission or clinical document needs stronger controls. Workflow sophistication should rise with impact: the higher the cost of a plausible but incorrect output, the more explicit the stages, evidence requirements and approval gates must become.
Feedback loops turn usage into performance
Many deployments collect activity metrics but no learning signal. Leaders know that 600 employees opened the assistant and generated 18,000 messages, yet cannot say whether proposals improved or service resolution accelerated. Adoption is not evidence of performance. A workflow needs feedback tied to its purpose: acceptance rates for generated code, edit distance for drafted text, escalation rates for support responses, forecast accuracy for sales recommendations or defect rates for extracted data.
Take a customer-service assistant that proposes replies. If agents routinely rewrite its answers, those edits are valuable operational data. Categorising the reasons—incorrect policy, unsuitable tone, missing account context or excessive length—reveals whether the remedy is better retrieval, revised instructions, cleaner source material or a different model. Merely asking agents whether they ‘liked’ the tool produces weaker evidence and encourages product teams to optimise sentiment rather than outcomes.
Feedback must also arrive quickly enough to influence the system. Quarterly reviews are too slow for workflows used thousands of times a week. Teams should inspect a representative sample regularly, compare AI-assisted and unassisted performance, and route recurring failures to an owner. Even a weekly review of 100 cases can expose systematic defects that aggregate dashboards conceal. The aim is not perfection; it is a controlled loop in which errors become design inputs rather than recurring surprises.
Human review must be designed, not assumed
‘Human in the loop’ is frequently treated as a universal safeguard. In practice, a rushed employee approving 80 plausible outputs is not meaningful oversight. Automation bias encourages reviewers to accept recommendations, especially when explanations are fluent and the source evidence is hidden. Review therefore needs structure: clear rejection criteria, access to supporting material, manageable volumes and authority to override the system without penalty.
Different risks demand different review patterns. A content team may sample 10 per cent of low-risk social posts while checking every regulated claim. A finance function might allow automatic coding of invoices below £500 when supplier, purchase order and amount all match, but require manual review for exceptions. A software team can permit AI-generated tests to run automatically while preventing generated code from merging until tests, security scans and peer review pass. These are workflow rules, not prompting techniques.
Organisations should also account for the cost of review. If AI saves five minutes of drafting but creates eight minutes of verification, the workflow has shifted labour rather than removed it. Verification can still be worthwhile when it improves quality, but the benefit must be explicit. Good design presents evidence beside the output, highlights uncertainty and directs attention to exceptions. The model should reduce the reviewer’s search burden, not merely generate more material for inspection.
Context quality is an operational responsibility
When an assistant gives inconsistent answers, the model is not always at fault. It may be drawing from duplicated policies, expired price lists or documents written for different regions. Retrieval-augmented systems expose weaknesses that organisations previously tolerated because experienced employees knew which files to ignore. AI makes knowledge debt visible at scale.
Reliable context requires ownership. Documents need dates, jurisdictions, permissions, version status and a named maintainer. Sensitive information must be separated from material that can safely enter a model’s context. Access controls should follow the user and the task; a broad search index that reveals payroll data to a general assistant is not an innovation failure but a governance failure. Teams also need rules for conflicts, such as preferring an approved policy over an older presentation that describes the same process differently.
More context is not automatically better. Loading dozens of documents can increase cost, slow responses and bury the decisive fact. A procurement assistant evaluating a contract renewal may need the contract, usage history, service incidents and current pricing, not the entire company drive. Context should be selected around the decision and tested using real cases. This requires collaboration among domain experts, data owners, security teams and product designers—precisely the cross-functional work that a prompt workshop tends to avoid.
Build teams around workflow ownership
AI programmes often divide responsibility poorly. Technology teams select models, business teams supply use cases, risk teams approve policies and employees are expected to connect the pieces. When performance disappoints, no one owns the end-to-end system. A better structure appoints a workflow owner accountable for the business outcome, supported by domain specialists, technical staff, operations and risk. The owner does not need to build the model, but must understand how every stage affects the decision.
Capability building should reflect this reality. Employees do need model literacy: they should understand hallucination, context limits, data sensitivity and the value of precise instructions. But training should also teach process mapping, evaluation, exception handling and escalation. Instead of asking learners to create the cleverest prompt, organisations should ask them to redesign a real task, define success, identify failure modes and run a controlled trial.
A disciplined pilot can be small. Choose one team, one decision and perhaps 200 historical cases. Establish a baseline for time, quality and error cost; test the assisted workflow; then compare results. If proposal preparation falls from four hours to two but factual corrections rise from one to six per document, the gain is not yet secure. That evidence directs improvement far better than anecdotes about impressive outputs.
The competitive advantage is operational
Foundation models are increasingly available to every serious organisation. Prompt patterns spread quickly, model capabilities converge and employees move between firms. Sustainable advantage is therefore unlikely to come from possessing a secret phrase. It will come from integrating AI into proprietary processes, trusted data, clear accountabilities and rapid feedback loops that competitors cannot copy overnight.
This changes the executive question. Rather than asking how many people have completed AI training, leaders should ask which decisions now happen faster or better, where error rates have changed, how exceptions are handled and who owns improvement. They should fund evaluation and knowledge maintenance alongside licences. A £500,000 platform deployed into chaotic processes can amplify confusion; a modest tool attached to a well-designed workflow can deliver measurable returns.
The AI skills gap is real, but it is routinely defined too narrowly. Organisations need people who can frame decisions, expose assumptions, design hand-offs, measure outcomes and adjust controls as evidence changes. Prompting belongs inside that discipline, not above it. The companies that learn this will treat AI adoption less like installing software and more like redesigning work—and their results will show the difference.
Comments (0)
Discussion is opening soon. Be the first to comment.