That sounds like a dodge, so the rest of this post makes it concrete: what deflection means, why published numbers are close to useless for comparison, how to calculate your own ceiling before you pay for anything, what raises the number, what quietly caps it, and what a small team should watch instead once the AI is live.
What does deflection actually mean, and why does every vendor define it differently?
Deflection has no shared definition across the support industry, which is the first reason quoted rates do not compare. Some tools count a conversation as deflected whenever the AI sent any reply at all, even if a teammate finished the job a minute later. Some count conversations that closed without a human message, which quietly includes everyone who got frustrated and left. Some count help center article views as deflections, on the theory that a read article is a ticket never filed. Others count only sessions where the visitor never clicked the button asking for a person.
For your own purposes, use the strictest definition you can measure, because it is the only one that predicts your workload. A conversation is deflected when the customer got a complete and correct answer, no teammate had to touch it, and the same customer did not come back about the same problem a day or two later. That last clause matters more than people expect. A bot that confidently gives a wrong shipping policy looks like a deflection on Monday and becomes an angry email on Wednesday, with the original ticket still counted as a success in the dashboard.
Why can't anyone give you a single honest number?
Because deflection is mostly a property of your ticket mix, not of the AI. Two businesses running identical software land in very different places depending on what customers ask.
An online store with steady questions about delivery times, returns windows, sizing and stock is sitting on a large pile of questions whose answers are stable, public and identical every time. A B2B tool whose inbox is full of "why did this integration stop syncing for my account" is sitting on a pile of questions that each need a lookup, a log, or a judgment call. The store will see a much higher share handled without a human, and neither business learned anything about the software from that gap.
Two more factors move it. The first is documentation quality: AI answers from what you have written, so a thin or contradictory knowledge base sets a hard ceiling no model can climb over. The second is traffic mix. Pre-sales visitors browsing your website ask broad, repeatable questions. Logged-in customers with a broken order ask specific ones. A site with heavy top-of-funnel traffic will show a flattering number that says more about its marketing than its support.
Small teams have one extra complication: sample size. When your weekly volume is modest, a single unusual week of incidents or a promotion can swing the measured rate hard in either direction. Judge the trend over a longer stretch and do not react to one quiet Tuesday.
How do you work out your own realistic rate before you buy?
This is the part almost nobody does, and it takes one afternoon. Export your last few hundred conversations and tag each one into a bucket:
- Answerable from information you have already published. Policies, hours, pricing, how something works, where to find a setting.
- Answerable only with account or order data. Correct answer exists, but it needs a lookup in another system.
- Requires an action on your side. A refund, a cancellation, a replacement, an exception to policy.
- Requires judgment or authority. Negotiations, complaints about your team, anything legal or regulated.
- The customer explicitly wants a person. Sometimes stated in the first message.
Bucket one is your realistic starting point. Bucket two becomes reachable once you connect the systems that hold the data, which is why an AI wired into your CRM answers questions one that is not simply cannot. Buckets three through five are where humans earn their keep, and trying to automate them is how teams get bad reviews.
Now discount bucket one twice. Take out the questions where your written answer is currently vague, outdated or buried, because the AI will inherit that vagueness. Then take out a slice for people who will ask for a human anyway. What remains is a defensible expectation you can hold a vendor to during a trial, and it is specific to your business rather than to someone else's case study. Re-run the exercise after a month or two; the tagging is faster the second time and the drift tells you what to write next.
What lifts the number, and what quietly caps it?
Three things lift it reliably. The first is writing your knowledge base as answers to questions rather than as topics, because that is the shape a customer's message arrives in. The second is a weekly habit of reviewing the questions the AI could not answer and publishing those answers; treat that log as a content backlog rather than a failure report, and the curve keeps climbing. A good knowledge base is the single highest-leverage thing a small team controls here.
The third is letting the AI complete tasks instead of only describing them. A customer who wanted to speak to sales and instead got a meeting booked on your calendar did not need a teammate. Corebee answers from your website and knowledge base, books meetings through Cal.com or Calendly, and syncs to HubSpot, Salesforce, Pipedrive, monday and Zoho, which turns a chunk of bucket two into questions that resolve without anyone reading them.
The caps are worth naming plainly. Anything requiring authority to bend a policy will escalate, and should. Anything ambiguous in your own policies will escalate, because the AI has nothing firm to stand on. Regulated advice should escalate by design. And handoff quality sets an invisible ceiling: if a customer asks for a human and gets stalled, word spreads in reviews and everyone starts typing "agent" as their opening message. An AI that hands off cleanly when asked or when unsure ends up deflecting more over time than one tuned to never let go.
Does the number look different on chat, email and phone?
Yes, and averaging them together hides the useful detail.
Web chat usually shows the highest share handled without a human. Messages are short, questions repeat, and much of the traffic is pre-sales. Email skews lower: threads are longer, often contain several questions at once, and lean account-specific. The realistic win on email is frequently a strong draft that a teammate approves in seconds rather than a fully untouched reply, which is worth tracking separately in your shared inbox.
Messaging channels sit in between. WhatsApp, SMS, Instagram, Facebook and Telegram bring informal, asynchronous, order-status-heavy traffic that automates well, as long as the same knowledge answers every channel instead of each one having its own half-configured bot. Voice is its own case. Callers often want confirmation more than information, and an AI receptionist that states it is an AI, answers the common questions, books the appointment and routes the rest is doing real work even when a human eventually picks up the thread.
How should pricing change what you count as a win?
This is where deflection stops being a support metric and becomes a purchasing one, because the pricing model decides whether a deflection saves you money or just moves it.
With per-resolution pricing, every automated answer carries a cost, so the incentive is to grow the count of things labeled resolved. With per-seat pricing, the cost sits on your team instead, and adding a weekend helper or giving a warehouse colleague access to answer one question a week starts to look expensive. Both models are perfectly workable; they simply mean your deflection target is entangled with your bill.
The major platforms sit at different points here and are strong for different reasons. Zendesk and Freshdesk are mature suites with deep routing, reporting and enterprise workflow, generally licensed per agent with AI capabilities layered on. Intercom is known for its AI agent and pairs it with usage-based charging for automated resolutions. Gorgias is built tightly around ecommerce operations. Help Scout keeps a simple, well-mannered shared inbox. Tidio focuses on small-business chat and sales. Chatwoot is open source, which is attractive if you would rather self-host and maintain it yourself. Pick on fit, not on whose deflection claim is loudest, and if you are weighing the incumbents, an honest comparison of Zendesk alternatives beats a feature grid.
Flat pricing removes the arithmetic entirely. Corebee is $99 per month with unlimited seats and unlimited AI conversations, no per-seat fee and no per-resolution fee, so a deflection is worth exactly what it saves your team in time and nothing is charged for trying. Setup takes about ten minutes and needs no developer, which matters when nobody on your team has spare hours to run an implementation project. The pricing page has the details.
What should a small team measure instead?
Deflection rate is a fine byproduct and a poor target. Four measures tell you more.
First, hours of repetitive replying returned to your team, which is the actual thing you bought. Second, after-hours coverage: the count of questions that got a real answer while nobody was awake, compared with what used to sit unread until morning. For a small team this is often the largest gain and it never shows up as a deflection improvement, because those conversations were previously just delay. Third, handoff accuracy, split two ways: escalations that should have been answered, and answers that should have been escalated. The second kind is the expensive one. Fourth, satisfaction on AI-handled conversations measured against human-handled ones, watching for a widening gap rather than an absolute score.
If those four are moving the right way, your deflection rate is healthy whatever it happens to read, and if they are not, a high deflection rate is only telling you how many customers gave up quietly.