The Magic Quadrant won't tell you if your bot is breaking the law

Gartner published its 2026 Magic Quadrant for Conversational AI last week, and the industry did what it always does: argued about the boxes.
The complaints are not unreasonable. Wayne Butterfield laid out a good list of them on LinkedIn - NiCE Cognigy dropped from Leader to Visionary apparently for the crime of being acquired and better integrated; boost.ai marked down despite market-leading customer satisfaction; Rasa, Parloa and Decagon not in the quadrant at all despite some of the largest valuations in the space; Salesforce somehow a Leader. His conclusion is the honest one, and it's the one I want to build on:
I do find the Magic Quadrants a great discussion point, and debating tool, but I'd never use them to actually choose a vendor that's right for me.
Right. So let's talk about the thing the quadrant can't put on either axis, and which matters more than any of it.
TL;DR
- The MQ ranks vendors. Your regulatory exposure is a property of your deployment, and it doesn't move when your vendor changes quadrant.
- Every tool on that chart is getting better at sounding human. That is precisely the capability regulators have written law about.
- "Ability to Execute" and "Completeness of Vision" contain no axis for "will disclose it's an AI when a real caller asks" - and that gap is where the fines live.
- You cannot buy your way out of this by picking a Leader. The obligation is yours the moment the call connects. What you can do is verify it, continuously, from the outside.
The quadrant measures the vendor. The regulator measures you.
Here's the category error underneath the whole debate.
A Magic Quadrant is a judgement about companies. Ability to execute, completeness of vision, market understanding, roadmap. Useful inputs to a procurement shortlist. Genuinely.
But not one of those axes is a statement about whether the bot you deploy on your number will, in production, on a Tuesday, tell an 82-year-old caller that it's an AI before it starts processing their renewal. That's not a fact about Google or Salesforce or Kore.ai. It's a fact about your prompt, your config, your guardrails, your last release - and it changes every time any of those do.
Gartner cannot measure it because it isn't the vendor's to own. Two firms can deploy the exact same Leader-quadrant platform and one is compliant and one is exposed, because compliance lives in how you built and configured it. This is the same reason I keep flipping the testing pyramid: the interesting failures are at the boundary between the tool and your specific use of it, and no vendor scorecard reaches that boundary.
So when the debate is "is SoundHound really above boost.ai," the more expensive question goes unasked: does the thing I actually shipped do what the law requires? The quadrant is silent on that for every vendor on it, Leader and Niche Player alike.
"Presents as human" is a feature on the roadmap and a liability in the call
Read what everyone on that chart is racing toward. More natural turn-taking. Sub-second latency. Barge-in. Emotional prosody. Voices you can't clock as synthetic. It's genuinely impressive engineering, and every vendor is measured partly on how far along it is.
Now read Article 50 of the EU AI Act: an AI system that interacts with a natural person must inform them they're dealing with an AI, unless it's obvious from context. The better your vendor gets at the thing the quadrant rewards - sounding human - the less obvious it is from context, and the more explicit your disclosure obligation becomes.
Sit with that. The exact capability that moves a vendor up and to the right is the exact capability that increases your regulatory duty. The chart rewards the thing that raises your risk, and says nothing about the mitigation. A more human-sounding bot is a better product and a bigger liability in the same breath, and only one of those shows up in the quadrant.
And it's not one rule. A modern voice channel sits inside something like thirteen overlapping regimes - EU AI Act Articles 5 and 50, FCA Consumer Duty, FG21/1 on vulnerable customers, Ofcom's General Conditions, UK and EU GDPR on biometric voice and emotion data, PECR, the Equality Act, the Unfair Commercial Practices Directive, and more. None of them were designed together. None of them care which quadrant your vendor is in. All of them apply to the call your customer just had.
Where a scorecard sends you wrong
The practical danger isn't that the MQ is imperfect. It's that it invites you to treat vendor selection as the compliance decision. Pick a Leader, tick the box, move on. That instinct is exactly backwards for three reasons.
It assumes compliance is procured, not operated. Most teams treat regulatory conformance as something they did at procurement - a vendor questionnaire, a signature, a closed deal. Nobody is continuously checking that the bot in production still behaves the way the questionnaire promised. The honest version of "are we compliant?" for most teams is "we hope so, and nobody has complained yet."
It assumes the vendor's behaviour is your behaviour. It isn't. The disclosure, the consent handling, the vulnerable-caller routing, the recording notice - those are emergent from your configuration and your prompts. The platform gives you the means. Whether you actually did it is a separate, testable fact.
It assumes stability. A quadrant is an annual snapshot. Your bot changed this morning. Someone tuned a prompt, swapped a voice, added an intent. A voice channel is a living thing, and a once-a-year judgement of your vendor tells you nothing about the version answering calls right now.
What actually answers the question
The only way to know whether your voice AI meets its obligations is to have something that isn't your voice AI call it up and check - against the actual text of the rules, not a vendor's marketing.
That's what we built. You point an audit at your real number, pick the regulations that apply to your business, and a synthetic caller runs the catalogue: does it announce it's an AI in the first turn, does it disclose recording and data use, does it slow down for a confused caller, does it take "uh-huh, sure" as consent for a policy renewal? We ran exactly this against a live UK number and ten of eighteen tests failed, six of them critical. Every verdict quotes the line from the transcript and the rule it breaches, so your auditor can check the working.
Three things follow, and they're the three the quadrant can't give you.
It measures your deployment, not your vendor. The verdict comes from a call placed at your number, graded against criteria you chose. Vendor-neutral by construction - we test over real PSTN, so whatever's in the top-right box, if it terminates a phone number we can audit it.
It runs on a schedule, not once a year. Your bot drifts continuously; the audit does too. Compliance becomes something you run like a test suite, not a £15k report that's out of date the day it lands.
It's evidence, not opinion. A dated, machine-readable record that the journey did the right thing on the day it mattered. The kind of thing the FCA asks for, that a quadrant screenshot won't stand in for.
By all means enjoy the quadrant
I'm not here to tell you to ignore it. Butterfield's right that it's a great discussion point and a fun thing to argue about, and there's real signal in the movements if you read them carefully. Debate whether Salesforce earned its box. Enjoy it.
Just don't mistake it for the answer to a question it never asked. The quadrant tells you which vendors Gartner rates. It does not tell you whether the bot on your number would survive a regulator listening to one of its calls. Those are different questions, and only one of them ends in a fine.
Pick whichever vendor fits. Then find out - from outside the system, against the actual law - whether what you shipped is legal.
Frequently asked questions
Doesn't picking a top-quadrant vendor reduce my compliance risk? Marginally, in that mature platforms give you better controls. But the obligation is on you as the deployer, and two firms on the same platform can land on opposite sides of a compliance line. The tool is necessary, not sufficient.
We already did a compliance sign-off at procurement. That was a point-in-time claim about a bot that has since changed. Procurement conformance and production conformance are different facts. The second one is the one a regulator samples.
Which regulations does this cover? The catalogue maps to EU AI Act Articles 5 and 50, FCA Consumer Duty and FG21/1, Ofcom General Conditions, UK/EU GDPR, PECR, the Equality Act and more - sector-weighted, so financial-services workspaces see PRIN 2A near the top and telco workspaces see Ofcom conditions first.
Does it matter which platform we're on - Genesys, Amazon Connect, a CAI vendor from the quadrant? No. We test over real PSTN, so anything that terminates a phone number is testable, vendor-neutral.
If you've deployed one of these very impressive, very human-sounding bots and you'd like to know whether it actually meets its obligations - not whether its vendor is a Leader - that's the job we built TotalPath for.
Or register for a free account and run one against your own number this afternoon.
Want us to explore your IVR?
TotalPath runs the same kind of test against your stack. Real audio. Real findings.