Why is "what percentage can AI handle?" the wrong question?
Key Takeaways
For busy support leads: the reason the percentage framing is so popular is that it is easy to put in a slide, and the reason it is useless is that it tells you nothing about which conversation in front of you right now should go where. Swap it for a two-axis sort you can complete in an afternoon with a list of your top topics and a whiteboard. The output is a routing policy you can defend to your CEO, rather than a target you inherited from a vendor's marketing page.
- 1Cost of being wrong beats complexity. A simple question with an expensive wrong answer should not be automated, however simple it is.
- 2Reversibility is the second axis. If a mistake can be undone in one click, automate it. If it moves money or data permanently, do not.
- 3The percentage is a result. It falls out of your topic sort. Setting it as a target inverts cause and effect.
- 4Draft and review is underrated. For the expensive-but-unemotional middle, a human sending an AI draft beats both extremes.
- 5Measure the same things for both. Different metrics for AI and humans hide the comparison you actually need to make.
Because it is an outcome of decisions you have not made yet, and treating it as an input produces bad decisions in both directions. Set a target that is too high and someone tightens the escalation rules to hit it, which means customers who need a person stop getting one. Set it too low and you keep humans answering questions your documentation already answers perfectly well.
It is also unknowable in advance for your specific business. The share of automatable conversations depends on your product's surface area, how well your content is written, whether your buyer is also your user, and how much your customers trust software with their problems. Two companies in the same category can differ by a factor of two on the same product. A published figure cannot know any of that about you.
Where research does exist, it measures productivity rather than a handling share, and the effect is uneven in a way an aggregate percentage would hide. A study of 5,179 support agents at one software firm found that access to an AI assistant "increases productivity, as measured by issues resolved per hour, by 14% on average, including a 34% improvement for novice and low-skilled workers but with minimal impact on experienced and highly skilled workers" (Brynjolfsson, Li and Raymond, NBER, 2023). Read that as a statement about who benefits, not about what share of your inbox is automatable.
There is a more practical objection. The percentage does not help an agent, a designer, or a router make a single actual decision. Given a specific conversation about a duplicate charge, "AI handles 60 percent" tells you nothing. "Anything that moves money over $50 goes to a person" tells you exactly what to do. Build policies at the level where decisions get made.
What should actually decide the routing?
Two questions per topic, both answerable in about thirty seconds.
What does a wrong answer cost? Count the direct cost, the time to unwind it, and whether the customer would reasonably tell other people about it.
How reversible is the outcome? Can the customer undo it themselves, can you undo it, or has something permanent happened such as money leaving, data being deleted, or a commitment being made on the record?
Those two axes give you four quadrants and a clear default for each.
| Cheap to be wrong | Expensive to be wrong | |
|---|---|---|
| Easily reversible | Automate fully | Draft and review, or automate with a visible person option |
| Hard to reverse | Automate, but confirm before acting | Human only |
Complexity does not appear anywhere in that matrix, and that is the point. A complicated multi-step troubleshooting walkthrough with a cheap, reversible failure mode is an excellent automation candidate. A one-line question about whether a customer's data was included in a security incident is trivially simple and should never be automated. Complexity is a proxy that fails exactly where the stakes are highest.
How do you sort your topics in an afternoon?
List your top twenty topics from actual conversation data, not from memory. For each one, write two numbers: the rough dollar cost of a wrong answer, and a one-to-three score for reversibility where one means the customer fixes it themselves and three means something permanent happened.
Then place each on the matrix and write the default next to it. That is your routing policy, and it will be more specific and more defensible than anything derived from a percentage.
| Topic (illustrative) | Cost if wrong | Reversibility | Default |
|---|---|---|---|
| How do I reset my password | Near zero | 1 | Automate |
| Which plan includes the API | Low, customer checks before buying | 1 | Automate |
| Why was I charged twice | Refund plus agent time | 2 | Draft and review |
| Can I get a refund outside policy | Money leaves, sets precedent | 3 | Human only |
| Was my data in the incident | Trust, potentially legal | 3 | Human only |
| How do I export my data | Low | 1 | Automate |
| Cancel my account | Revenue, hard to undo the relationship | 3 | Human only |
| Does this integrate with X | Low, but a wrong yes costs a churned trial | 2 | Automate with a citation |
The last row deserves attention. Integration and capability questions look harmless and are not, because a confidently wrong yes produces a customer who buys, discovers the gap, and churns angry. Automating them is fine, but only with the source cited so the customer can check, which turns an assertion into a verifiable claim.
What does a wrong answer actually cost?
Work an example so the axis stops being abstract. Illustrative inputs throughout.
A customer asks whether a return is still within the window. The assistant reads the policy wrong and says yes when it should have said no. You now either honour it, which costs the item value plus shipping, or refuse it, which costs an unhappy customer and about 20 minutes of agent time on a thread that is now adversarial. Call it $40 plus 20 minutes. At a loaded rate of $32 an hour, that is roughly $51 for one wrong answer.
If that topic appears 200 times a month and the assistant gets it wrong three percent of the time, that is six wrong answers, or about $306 a month. Now compare against the human cost of handling all 200: at four minutes each, that is 13.3 hours, or about $427 a month.
So automating this topic saves roughly $121 a month, before you account for the customers who never contacted you at all because they got an instant answer. That is a genuine but modest win, and it is the sort of calculation that changes decisions.
Now change one input. Make the topic a billing dispute where a wrong answer costs $300 and a complaint. Six wrong answers a month is $1,800, and the humans would have cost $427. Same simplicity, same volume, opposite conclusion. The error rate did not change. The cost of the error did.
Which conversations should never touch a bot?
Four categories, and we would hold these even at a large volume cost.
Anything about a security incident or account compromise. The stakes are asymmetric, the customer is frightened, and a wrong reassurance is far worse than a slow correct answer.
Anything where a person has told you they are leaving. A cancellation conversation is your last chance to learn something true, and an assistant will handle it politely and learn nothing.
Anything involving a commitment you would be held to. Custom pricing, contractual terms, dates you will be judged against. An assistant that promises a delivery date has made a promise on your behalf.
Anything where the customer has already asked for a person. This is not a category of topic, it is a category of moment, and it overrides everything else in this article.
Does draft and review beat full automation?
For the expensive-but-unemotional middle, often yes, and it is the most underused pattern in the category. The assistant retrieves and drafts, a human reads it, adjusts, and sends. You get the retrieval speed and the judgment together.
Two honest costs. It does not scale the way full automation does, because a human still reads every message, so it does not solve a volume problem. And review quality decays with familiarity: after a few hundred good drafts, people start skimming and approving, which is precisely when the bad one goes out. If you use this pattern, sample and audit sent drafts periodically rather than trusting that review is happening.
Where it is genuinely excellent is onboarding. A new agent reviewing and editing drafts learns your product, tone and policy faster than any documentation achieves, because they are seeing the correct answer next to the real question hundreds of times. Several teams we have talked to use it as a training phase rather than a permanent mode, which we think is the right instinct.
How should you measure AI and humans against each other?
On the same metrics, on comparable conversations, and with an explicit acknowledgement that the comparison is unfair in a specific direction.
The unfairness matters. If automation handles your easy topics, the human queue is left with harder ones, so a straight satisfaction comparison flatters the assistant. Compare within a topic, not across your whole inbox, or the number means nothing.
| What to measure | Why |
|---|---|
| Satisfaction, per topic | Cross-inbox comparison is distorted by topic mix |
| Resolution without a follow-up thread | Catches answers that looked fine and were not |
| Reopen rate within seven days | The single best proxy for a wrong answer that was accepted |
| Escalation rate, per topic | Shows which topics are miscategorised in your matrix |
| Wrong-answer cost, sampled monthly | Read 20 automated threads and price the mistakes |
That last row is the only one that connects measurement back to the decision framework, and it is the one nobody does. Twenty threads a month read by a human, with mistakes priced rather than counted, will tell you more about whether your routing is correct than any dashboard.
Should you tell customers they are talking to AI?
Yes, and in the EU you no longer have a choice. Article 50 of the EU AI Act became applicable on 2 August 2026 and requires that people interacting directly with an AI system are informed that they are, unless it is obvious (EU AI Act, Article 50). Article 99(4)(g) sets fines of up to 15 000 000 EUR or 3 percent of total worldwide annual turnover, whichever is higher (EU AI Act, Article 99). Obligations elsewhere are moving in the same direction, so check the current position with someone qualified rather than assuming your existing label is sufficient.
Setting the legal question aside, the practical argument is strong on its own. In a SurveyMonkey non-probability online panel of 2,017 US adults fielded in December 2025, 84% said human agents are more accurate than AI and 81% said they believe AI is used primarily to save money rather than to improve service. Those are attitudes measured by a vendor on a non-probability panel, not facts about your product, but they describe the assumption your assistant walks into. Customers who know they are talking to software calibrate their expectations and ask cleaner questions. Customers who discover it later feel deceived, and that feeling attaches to your company rather than to the software. The downside of disclosure is small and the downside of being caught is large.
What matters more than the label is the exit. A clearly labelled assistant with an obvious route to a person is fine. A clearly labelled assistant with no exit is worse than an unlabelled one, because you have told the customer exactly what is trapping them. The mechanics of doing this well are covered in our handoff guide.
Where this breaks
The matrix assumes your topics are separable and your costs are estimable. Several situations break one or both.
If your product is genuinely bespoke, agency and consulting work being the obvious cases, topics do not recur enough to sort. Route by account and skip this exercise entirely.
If you have very low volume, perhaps under a hundred conversations a month, the sorting effort exceeds the saving. Answer everything yourself, and revisit when the queue starts hurting.
If your customers are enterprise buyers with named contacts, the relationship is the product and cost-of-error analysis understates the damage from an impersonal response. Weight everything toward humans regardless of what the matrix says.
And if your knowledge base is poor, none of this applies yet, because the automated answers will be wrong in the cheap quadrant too. Fix the content first. Routing policy applied to bad content just distributes wrong answers more efficiently.
Where pricing quietly changes this decision
This is the part that rarely appears in framework articles and it changes real decisions. If your platform bills per automated resolution, every routing choice is also a spending choice, and improving your content raises your bill because resolutions are the billable unit. Teams under that model start making routing decisions for cost reasons rather than for customer reasons, which is a bad way to design a support experience.
Flat pricing removes that pressure. Corebee is $99 a month regardless of volume, so the only question left is whether automating a topic serves the customer. We think that is the correct question and the pricing model should not be arguing with it. If you want the reasoning laid out, our pricing page is one number and a short explanation.
The honest caveat: flat pricing is not a reason to automate more. It just removes a reason to automate for the wrong motive. The matrix above is what should decide, and it does not care who bills you how.
What should you change this week?
Pull your top twenty topics from real data. Score each on cost of being wrong and reversibility, and place them on the matrix. Expect two or three surprises, usually simple topics you had automated that turn out to sit in the expensive column.
Then sample twenty automated threads, read them yourself, and price the mistakes rather than counting them. That single hour will tell you more than a quarter of dashboard watching. If you want to test the sort against real conversations before committing, you can start a free trial and route a couple of topics at a time.