Picture an organisation doing everything right. An NPS programme is running. CSAT is measured quarterly. Complaints are logged. Six months ago it added a sentiment analysis module. Over the past year it collected more than 40,000 customer signals.
Now ask the board one question: which business decisions did customer data change last year? Not confirmed — changed. With an owner, a date and a budget line.
In most organisations the answer is silence. And that is not a measurement problem. Which is precisely why another measurement tool will not fix it.
Is this for you?
Three questions. If two answers are yes, the rest of this is about your situation.
- Do you have a running feedback programme (NPS, CSAT, surveys)?
- Are you under executive pressure to deploy AI?
- Do you not know how many decisions customer data changed last quarter?
If all three are no and you have no feedback programme yet, AI is not your first step. Start with the listening fundamentals.
Why listening grows and action does not
Forrester's 2025 global survey of VoC and CX measurement practitioners shows where the chain actually breaks. Only 27% effectively get insights to decision-makers in time. Only about half are confident in their ability to analyse what they already collect. Only half can connect CX metrics to business outcomes, and fewer than a third can set realistic targets. Root cause analysis is rare. Surveys dominate, while emails and conversations are seldom used well.
Meanwhile, in Gartner's October 2025 survey, 91% of customer service leaders reported pressure from executive leadership to implement AI.
Put those two facts together. The pressure to buy AI is close to universal. The ability to turn customer data into decisions is not. The outcome is predictable: organisations buy what demos well, not what closes the gap.
AI makes signals cheaper to produce. It does not make decisions cheaper to make. A tool bought without that distinction widens the gap rather than closing it.
The signal-to-decision gap
Definition
The signal-to-decision gap is the distance between the volume of signals an organisation collects about its customers and the number of decisions those signals actually changed.
Listening programmes measure the left side of that equation. Value is created on the right.
You can measure it in an afternoon. Three numbers for the last quarter:
- Signal volume. Feedback, complaints, chats, call recordings — everything collected.
- Decision count. How many decisions were made or changed because of those signals. A decision counts only if it has an owner, a date and a consequence: a changed process, a reallocated budget, a cancelled or corrected initiative.
- Cycle time. Average time from signal to decision.
The ratio: decisions per 1,000 signals. In the example above — 40,000 signals, three decisions — that is 0.075. In practice this number rarely exceeds one, and most organisations have never calculated it at all.
A useful rule: if you do not know your ratio, it is too early to buy an AI tool. Without it you have no baseline to compare the investment against in twelve months' time — and the vendor knows this perfectly well.
Five depth levels: what you are actually paying for
Almost every AI CX tool is sold with the same vocabulary. The difference shows up when you ask what the system does with a signal.
| Level | What the system does | What it means if you already have VoC |
|---|---|---|
| 1. Collection | Gathers feedback, surveys, recordings | You have this. You would be paying twice. |
| 2. Classification | Assigns topics, sentiment, categories | A commodity. The price falls every year and it creates no advantage. |
| 3. Interpretation | Identifies cause, journey context and size of impact | Value starts here. Few vendors volunteer a demonstration of it. |
| 4. Recommendation | Proposes a specific action, an owner and a priority | Directly reduces the signal-to-decision gap. |
| 5. Action | Initiates the task in an operational system itself | Powerful, but risky without governance. Requires a clear line of accountability. |
The rule is simple: if you already run a working VoC programme, you pay only for level 3 and above. Everything below it is either a feature of your existing platform or a commodity whose price is falling.
Static or intelligent: one test
The distinction is not "does it have AI". It fits in a single line:
A static tool answers the question you asked. An intelligent one shows you what you didn't ask.
A survey platform is static by construction. You see only what you put in the questionnaire, and only from those who agreed to answer — typically 2–10% of the base. The customer who left quietly is not in your questionnaire. That is not a flaw in the tool; it is its design. The flaw appears when a static tool is presented as a customer experience measurement system.
The test for a vendor: "Show me a theme the client wasn't looking for, that your system found — and what changed as a result." If the answer is about more accurate assignment of categories you already defined, that is level 2. If the system surfaced a problem nobody had articulated, that is level 3, and worth paying for.
The market map: four layers, two different purchases
"Which contact centre platform should we choose" is in practice two questions bought as one. This is where the expensive mistakes come from: buying pipes and expecting brains.
| Layer | What it actually solves | Who is in the market | Depth |
|---|---|---|---|
| 1. Channels (CCaaS) | Calls, queues, routing, agent desktop, channel consolidation | In Gartner's 2025 CCaaS evaluation (8 September 2025, nine vendors) the Leaders were NiCE, Genesys, AWS Amazon Connect, Five9 and Talkdesk; Content Guru was placed as a Challenger; Cisco, Zoom and Vonage as Niche Players; no Visionaries. Not evaluated but significant in the market: 8x8, Google, Microsoft, Sprinklr. | 1–2, up to 4 with add-ons |
| 2. Conversation intelligence | Analysis of every conversation rather than 3%: causes, quality, real-time agent assist | CallMiner, Verint, Calabrio, NICE Enlighten, Observe.AI, Cresta. This is where level 3–4 value actually lives. | 2–4 |
| 3. Feedback (XM / VoC) | Surveys, NPS, CSAT, closed loop | In Forrester's Q4 2024 evaluation of customer feedback management solutions (nine vendors, 26 criteria) the Leaders were Medallia and Qualtrics; InMoment, Sprinklr and Concentrix were also evaluated. | 1–3 |
| 4. Journey management | Journey metrics, orchestration, triggering action | Genesys, Medallia, Qualtrics and specialist journey tools. The youngest layer and the one carrying the heaviest promises. | 3–4 |
The practical conclusion: layer 1 does not close the signal-to-decision gap. It solves availability. Buying CCaaS in the hope of customer insight is paying for plumbing and expecting water quality analysis.
Four market moves worth knowing before you negotiate
- Verint and Calabrio are already merged. The deal closed on 26 November 2025 — the two largest independent workforce engagement and quality vendors became one. Fewer alternatives, and negotiating power shifts toward the vendor.
- Genesys took a $1.5bn investment from Salesforce and ServiceNow. CCaaS and CRM are converging. In 2026 you are not only buying a contact centre — you are buying a position in that war.
- Talkdesk Embedded puts contact centre components inside someone else's CRM. Platform boundaries are dissolving, and the "one platform for everything" argument is weakening.
- Zoom entered the evaluation; 8x8 dropped out of it. This layer moves faster than a five-year contract lasts.
Which yields a simple negotiating rule: commit long-term to infrastructure, never to intelligence. The intelligence layer will be unrecognisable in three years, and you will be tied to a 2023 vision of it.
And the mid-market reality
None of the leaders above were designed for a 25-agent contact centre in Vilnius, Tallinn or Porto. The gap between their pricing and implementation timelines and a mid-sized European budget is measured in orders of magnitude.
But something real has changed: the intelligence layer can now be decoupled from the platform. Transcription plus a language model via API, running on your existing telephony, delivers level 3 analysis for a fraction of an enterprise licence — including in smaller languages, which three years ago was the blocking constraint. It is not free: it needs an owner inside the organisation, a data processing agreement and quality control. But it is a genuine option that did not exist two years ago, and no vendor will propose it to you, because there is nothing in it to sell.
Six filters not every vendor passes
1. Proof of depth, not a demo
Ask: "Show me three recommendations your system produced for a real client last quarter, and what was done about them." If the answer is another dashboard screenshot, it is level 2, whatever it is called.
2. Signal coverage beyond surveys
The largest untapped value sits where the data already exists and nobody reads it: calls, chats, emails, complaint descriptions, free-text CRM fields. A tool that only analyses survey comments processes the smallest and most biased slice of your data — the customers who still agree to answer.
3. Accuracy in your language, not "support" for it
"We support 100 languages" is not an answer. Support is not accuracy. Ask for a number: what is topic assignment accuracy in your language, measured on what data, against what benchmark. Vendors trained predominantly on English data can drop sharply in smaller languages, and no one volunteers that in a demo. If the vendor has no such number, it does not exist — and you will have to produce it yourself.
4. The path to the decision owner
The question is not "is there an integration" but "in whose task queue does the conclusion land". An insight that travels to a dashboard opened once a month by a CX manager with neither budget nor authority to change a process will never become a decision. That is a governance question, not a technology one — and it has to be answered before the purchase, not after.
5. The legal boundary: the EU AI Act
a) The prohibition — Article 5(1)(f), in force since 2 February 2025. The EU prohibits AI systems that infer the emotions of workers in the workplace on the basis of biometric data — voice tone, facial expression — outside narrow medical and safety exceptions. Penalties reach €35 million or 7% of global annual turnover.
Where the line falls in practice: recognising a customer's emotion during a call is explicitly identified as not prohibited in the European Commission's guidelines. Assessing an employee's emotional state from their voice is prohibited. In many contact centre platforms both run inside the same module, and the employee-facing part is sometimes on by default.
b) Transparency — Article 50, applying from 2 August 2026. Businesses must clearly inform users that they are interacting with an AI system or receiving AI-generated content, and that information must be easily accessible. In a contact centre this means every automated conversation, email or chatbot response will need a clear AI marking. Penalties reach €15 million or 3% of global annual turnover. For generative systems already on the market before that date, the machine-readable marking duty under Article 50(2) was deferred to 2 December 2026 under the Omnibus provisional agreement of 7 May 2026.
Ask in writing: does the product infer employees' emotional state from voice, video or physiological data; is it enabled by default in EU deployments; can it be switched off at tenant level. An evasive answer is already an answer.
6. The cost of leaving
Who owns the topic taxonomy and the historical assignments? If the vendor does, then in two years you will be unable to compare your own data against anyone else, and your negotiating position at renewal will be zero. Negotiate export terms in the first contract, not the third.
Five questions worth sending before the demo
Not during it — before. In writing. The quality of the answers filters better than an hour of slides.
- How did you measure accuracy in our language — on what data, and against what benchmark? The likely answer is "we haven't". That is your information. Don't demand a third-party audit report; almost none exist in this market, and asking for a non-existent document gets you "we'll come back to you".
- Show three cases where the system recommended an action the client had not articulated — and what was done about them.
- Who owns the taxonomy and historical assignments after termination? Can we export them as CSV?
- Confirm in writing: does the product infer employees' emotional state from voice, video or physiological data? Is it on by default in EU deployments? Can it be disabled?
- Can a recommendation create a task in our operational system without a human in between — and who is accountable if the action is wrong?
A thirty-day pilot before you sign
A demo on vendor data proves nothing. The only serious test is your data and a benchmark the vendor cannot see.
Experiment structure
- Hypothesis
- We believe the vendor's model assigns topics and causes to our customer signals accurately enough to base operational decisions on its output.
- Change
- We will take 500 real records from our own channels (not demo data). Two independent colleagues will code them by hand — that is the benchmark. The vendor will not see the benchmark and will process the same 500 records through their system.
- Expected outcome
- Agreement with human coding of at least 80% at topic level and 60% at cause level. In addition, three of the system's recommendations go to a business owner with one question: "Would you have done anything about this?"
- Success metric
- Accuracy against the human benchmark, plus at least one recommendation the owner is prepared to implement without further analysis.
- Review date
- 30 days from the start. The decision is made the same day: continue or stop.
What this costs your team
The first question you will get from your executive team. Calculate it like this:
| Work | Who | Effort |
|---|---|---|
| Benchmark coding | 2 people | 500 records. Around 60–100 records per hour for short text, 20–30 for call transcripts. Only a 100-record sample is double-coded: enough to measure inter-rater agreement, and half the cost of coding everything twice. |
| Coordination | 1 person | ~1 day: data extraction, handover, coding rules |
| Analysis and decision | 1 person | ~1 day: comparison against the benchmark, review of three recommendations with the business owner |
For the voice channel that is roughly 6–8 working days. For text, less. Compare it with the sum you are about to commit for three years. A vendor who refuses this test has just saved you a year.
What can burn
A realistic picture is worth more than a sales case. The five risks that materialise most often:
| Risk | Likelihood | What to do |
|---|---|---|
| Recommendation reaches someone without authority | Very high | Before purchase, name who signs off on decisions arising from each type of insight. This risk does not disappear after deployment — it only gets more expensive. |
| Recommendations not implemented due to budget | High | Agree in advance on the threshold below which a decision proceeds without a separate board submission. |
| Agents feel monitored and stop raising issues | High | Involve agents in the pilot; communicate clearly that the process is being analysed, not the person. This is a legal question as much as a cultural one — see filter 5. |
| Data quality: messy records | Medium | Clean the data before the pilot, not after. Otherwise you will be measuring the state of your CRM, not the accuracy of the model. |
| Vendor locks in your data | Medium | Export terms in the first contract. |
On calculating ROI before the pilot
The vendor's proposal will almost certainly contain a calculation that adds saved agent hours, avoided churn and NPS uplift into one impressive figure. Three things are worth knowing about it before you take it to the board.
First, the value of a retained customer is usually calculated on revenue rather than margin — your CFO will spot that in three seconds. Second, avoided churn is a counterfactual: without a control group it cannot be proven. Third, the lines frequently count the same euro twice, because saved time and retained customers often come from the same fixed process.
An ROI calculation before the pilot is a sales artefact, not a financial model. The real number arrives on day 31.
A defensible model rests on two things only: hours saved as measured in the pilot, multiplied by fully loaded cost of employment; and one specific fixed process — with a before-and-after volume, a control period, and margin rather than revenue. One proven process convinces more than three forecast ones.
The question worth asking before any other
How many decisions did a customer signal change in the last three months — and who signed them?
If the answer is fewer than three, your problem is not a measurement tool. It sits between insight and decision — and AI does not yet work there, because that stretch is organisational, not technical.
A good AI tool is still useful in that situation. It simply has to be bought as a way to accelerate decisions, not as another layer of listening. The difference is visible in the first demo — if you know what to ask.
The buying kit: gates, questions and pilot protocol
A printable worksheet for vendor meetings: four gates, the full question set with space for answers, the 30-day pilot protocol and a defensible ROI structure. Six pages.
Download the worksheet (PDF)No registration · No email · Just the file
Sources. Forrester's 2025 survey of VoC and CX measurement practitioners (311 respondents, March–April 2025; a self-selected sample, so the findings are descriptive rather than representative) — forrester.com. Gartner survey of 321 customer service and support leaders, conducted October 2025, published 18 February 2026 — gartner.com. Regulation (EU) 2024/1689 (AI Act), Article 5(1)(f) (applying from 2 February 2025) and Article 50 (applying from 2 August 2026; Article 50(2) for systems already on the market from 2 December 2026, per the Omnibus provisional agreement of 7 May 2026); European Commission guidelines on prohibited AI practices, C(2025) 884 final. Market map: Gartner Magic Quadrant for Contact Center as a Service, 8 September 2025 (nine vendors); The Forrester Wave™: Customer Feedback Management Solutions, Q4 2024 (nine vendors, 26 criteria).
Vendor positions reflect publicly reported analyst evaluations as at the date of those reports and change over time. Analyst firms do not endorse any vendor, and this list is not a recommendation. Prices are deliberately omitted: they are negotiated, frequently confidential, and publicly circulating figures for the same product differ by an order of magnitude.
This article is not legal advice. For an assessment of your specific case under the AI Act, consult a lawyer.