The deployment started with a narrower promise than the sales pitch
Six months ago, a 420-person B2B software company put an AI layer into production across sales operations. The business had 96 account executives, 38 sales development representatives and a revenue operations team of 14. Its objectives were deliberately practical: reduce manual CRM work, improve forecast hygiene, surface stalled opportunities and give managers earlier warning of pipeline risk. The system connected to Salesforce, Gong, Outlook, Slack and the company’s product-usage warehouse. It summarised calls, suggested field updates, scored deal risk and drafted weekly forecast commentary.
The implementation was not a clean-room experiment. The company sold three product lines across North America and Europe, used seven opportunity stages and had accumulated nine years of inconsistent CRM records. Regional teams interpreted qualification criteria differently. Renewals were mixed with new business in some dashboards but separated in others. Approximately 18 per cent of open opportunities had no recorded next step, while 27 per cent had a close date that had already passed. AI did not arrive to optimise an orderly machine; it arrived to operate inside the same imperfect system sales operations teams manage every day.
The initial business case projected 11,000 hours of annual administrative savings and a four-point improvement in forecast accuracy. After six months, the outcome was less dramatic but more credible. The company was on track to save roughly 6,800 hours a year, forecast variance had narrowed by 2.6 percentage points, and managers were spending less time assembling information. Yet several apparently straightforward automations failed, especially where the system had to infer commercial judgement rather than summarise evidence. The distinction between those two tasks shaped nearly every result.
Adoption depended on removing work, not showcasing intelligence
The first month produced impressive demonstrations and mediocre usage. Seventy-eight per cent of sellers tried the AI-generated deal summaries, but only 34 per cent used them again in the following week. Reps said the feature was accurate enough, yet it lived in a separate panel and required them to change their workflow. By contrast, automated call notes placed directly into the opportunity record reached 83 per cent weekly adoption by month three. The notes did not feel futuristic; they simply removed an unpopular task.
The adoption pattern exposed a common error in enterprise AI programmes: counting feature activation rather than sustained behaviour. Sales operations initially reported that 91 of 96 account executives had “adopted” the platform because they had opened it once. A stricter measure—using an AI-assisted action at least three times a week—put genuine adoption at 57 per cent. After the team embedded suggestions inside Salesforce, reduced notification volume and allowed one-click acceptance of field updates, that figure rose to 76 per cent by month six.
Mandates played a smaller role than local credibility. Two regional managers reviewed AI alerts during weekly pipeline meetings and challenged false positives in public. Their teams adopted the system faster because reps could see how recommendations affected real decisions. Another manager treated the risk score as an inspection tool, and usage fell as sellers withheld notes or dismissed suggestions defensively. The lesson was operational rather than cultural: AI adoption improved when the system returned time or sharpened a conversation, and deteriorated when it became another mechanism for surveillance.
Data quality improved only after the model exposed its cost
The deployment quickly found that most data-quality problems were not missing-data problems. They were definition problems. “Decision maker identified” meant a named economic buyer in one region, attendance by any director in another, and little more than a checked box elsewhere. The model interpreted the field literally and repeatedly marked weak deals as healthy. During the first eight weeks, 41 per cent of opportunities labelled low risk had no confirmed budget owner in call transcripts or emails.
Rather than retrain the model immediately, revenue operations rewrote five qualification fields, added evidence requirements and retired 23 redundant CRM properties. The AI then proposed updates using specific source material: a customer quotation, meeting date or email reference. Acceptance rates for suggested field changes rose from 46 per cent to 71 per cent. More importantly, disputed updates became auditable. A rep could reject “legal review started” because the customer had merely requested a security document, and that correction fed the next rules revision.
The clean-up had a measurable cost. Three operations analysts spent almost six weeks mapping fields, testing historical records and reconciling regional definitions. That work was absent from the original implementation estimate, which assumed existing CRM data was broadly usable. It was also one of the programme’s most valuable outcomes. By month six, overdue close dates had fallen from 27 per cent of open opportunities to 9 per cent, while opportunities without a next step dropped from 18 per cent to 6 per cent. AI did not fix data quality autonomously; it made inconsistency visible, frequent and expensive enough to address.
Forecasting became more disciplined, but not autonomous
Before deployment, the company’s monthly forecast finished an average of 9.8 per cent away from actual bookings. Six months later, the rolling average variance was 7.2 per cent. The gain came primarily from earlier detection of slippage. The model combined stage age, meeting activity, stakeholder coverage, procurement language and product engagement to identify deals whose close dates were implausible. In the final month, it flagged 37 opportunities worth £6.4 million; managers moved 21 out of the quarter before the formal forecast call.
That was useful, but the model remained weak around discontinuities. It missed a £480,000 expansion that accelerated after a competitor’s security incident, because historical patterns suggested a 90-day cycle. It also downgraded a £310,000 public-sector deal when communication paused during a procurement blackout, despite the pause being expected. AI was strongest when current behaviour resembled known deal patterns and weakest when external events changed the commercial context.
The company therefore kept human-owned forecast categories and used AI as a challenge layer. Managers received three outputs: an evidence-based risk score, the factors driving it and a comparison between the rep’s close date and the model’s expected range. They were not given an automated booking number to accept blindly. This design reduced arguments about intuition without pretending uncertainty had disappeared. Forecast calls became 18 minutes shorter on average, largely because teams discussed exceptions rather than reading every opportunity line by line.
Productivity gains were real, uneven and easy to overstate
The headline productivity result was a reduction of 42 minutes per rep per week in CRM administration, measured through activity logs and a time study involving 31 sellers. Call summaries saved the most time, followed by automated contact creation and drafted follow-up emails. Across the sales organisation, the annualised saving was approximately 6,800 hours. That was substantial, but 38 per cent below the business case because many suggested actions still needed review and because some saved time shifted into higher-quality account preparation rather than additional selling calls.
Output did not rise uniformly. Sales development representatives increased completed personalised outreach by 12 per cent after the system drafted research briefs and first-pass emails. Account executives showed only a 3 per cent increase in customer-facing activity. Their work contained more complex discovery, negotiation and internal coordination, where generic automation offered less leverage. The strongest performers used AI to prepare questions and map stakeholders; weaker performers often accepted bland drafts that required extensive editing or generated poor replies.
Management productivity improved more clearly. Automated weekly summaries reduced preparation for pipeline reviews from a median of 74 minutes to 29 minutes per manager. However, the company stopped using AI-generated coaching recommendations after finding that they overemphasised easily measured behaviours such as talk ratios and question counts. A rep asking fewer, sharper questions could be rated below a colleague mechanically following a script. The surviving productivity gains came from compression and retrieval, not automated judgement about selling quality.
Exception handling became the hidden operating model
The original workflow assumed the AI would update routine records while people handled ambiguous cases. In production, defining “ambiguous” became the core design problem. During month two, automated opportunity-stage changes produced 186 reversals. Common causes included customers discussing implementation before signing, partners using internal project language and renewal calls being mistaken for expansion opportunities. The company suspended automatic stage progression and replaced it with approval prompts.
A tiered exception model followed. Low-risk actions, such as attaching a call summary or creating a contact already present in an email thread, happened automatically. Medium-risk actions, including changing a next step or adding a competitor, required one-click approval. High-impact changes—forecast category, close date, deal value and stage—remained human controlled. This reduced weekly reversals to fewer than 20 while preserving most of the administrative saving. It also made accountability clear when a disputed change affected a forecast.
Exceptions needed staffing, not merely software. Revenue operations created a rotating duty covering failed synchronisations, duplicate contacts, disputed recommendations and policy questions. The queue averaged 63 items a week in month one and 24 by month six, requiring roughly eight staff hours weekly. That workload was manageable, but it disproved the assumption that AI automation would be maintenance-free. Production systems need thresholds, escalation routes and owners who can distinguish a model error from a process defect or an unusual but legitimate deal.
The automations that survived were narrow, visible and reversible
By the six-month review, five capabilities had become dependable: call summarisation, contact capture, next-step suggestions, stale-pipeline alerts and draft forecast commentary. Each had bounded inputs, a clear user, and an obvious way to correct mistakes. Together they accounted for nearly 80 per cent of measured time savings. More ambitious functions—automatic stage changes, coaching scores and fully generated account plans—were either withdrawn or restricted to pilots.
The economics also became clearer. Annual software and infrastructure costs were £238,000, while implementation and data work added £164,000 in the first year. Using a conservative loaded labour rate, the annualised time saving was valued at about £285,000. Reduced forecast preparation and avoided pipeline clean-up added an estimated £96,000, putting first-year payback close to break-even rather than the four months originally forecast. The second-year case looked stronger, provided adoption held and integration costs did not climb.
The durable result was not an autonomous sales operation. It was a better instrumented one. Reps did less transcription, managers received earlier warnings, and operations teams gained evidence about where definitions and workflows failed. The company’s next phase will focus on renewal risk and territory planning, but with a stricter test: every automation must remove a specific task, cite its evidence, expose uncertainty and offer a quick reversal. Six months in production showed that AI created value when it respected operational boundaries. Where it tried to replace commercial judgement, reality imposed them anyway.
Comments (0)
Discussion is opening soon. Be the first to comment.