What counts as a working AI support agent?
Vendors tend to define success as containment: the share of conversations that never reach a human. That is a useful operational signal, but on its own it rewards the wrong behavior, because an agent that stonewalls a customer until they give up looks identical on a dashboard to one that solved the problem. A better definition has three parts. The customer got an accurate answer or a completed action. They did not have to ask again through another channel. And your team did not spend time repairing the interaction afterward.
Write that definition down before you look at a single report, because every later argument about metrics is really an argument about this definition. Owners and support leads who skip this step end up with a screen full of encouraging numbers and a nagging sense that customers are less happy than they were. A practical test: take your five most common questions and describe, in one sentence each, what a good outcome looks like. If you cannot do that, no metric will rescue you, and no vendor's default report will do the thinking on your behalf.
Which numbers tell the truth, and which ones flatter the tool?
Flattering metrics share a trait. They count activity rather than outcomes. Total conversations handled, messages sent, average first response time, and raw deflection all rise when the agent is fast, chatty, and unhelpful. Response time becomes close to meaningless once an AI is answering, since every AI reply is fast by construction. Treat it as a health check on your infrastructure, not as evidence of quality.
Metrics that resist flattery are harder to game because they depend on the customer coming back, or not coming back. Full resolution rate is the anchor: did the same person contact you again about the same topic within a sensible window? Around it, put escalation rate broken down by reason, repeat contact rate, and sentiment collected specifically on AI-handled threads rather than blended with human ones. Blending is the most common reporting mistake, because strong human performance masks weak automated performance and you lose the ability to see either clearly.
Then add one qualitative measure that no dashboard produces: a regular sample of transcripts read by someone who knows the business. Numbers tell you where to look. Transcripts tell you what is actually going wrong, and they are usually the fastest route to a fix.
How do you measure resolution without fooling yourself?
Resolution is the metric most worth getting right and the easiest to distort. Start by deciding what the same issue means. A customer who asks about a delivery on chat, then emails about the same order the next morning, represents one unresolved issue, not two separate contacts. Counting them separately makes your agent look busier and better at the same time, which is precisely backwards. That means conversations from every channel need to land in one place against a shared customer record, which is why a unified inbox matters more for honest measurement than it does for day to day convenience.
Next, choose a follow-up window that matches your product. A software subscription question is probably resolved if nobody writes back within a few days. A question about a shipment in transit may need longer. Pick the window deliberately, keep it fixed, and note it next to the number so nobody quietly shortens it later.
Finally, split resolutions into two buckets: questions answered, and actions completed. Telling a customer your refund policy is not the same as issuing the refund, and booking a meeting is not the same as describing your availability. Lumping the two together hides the most important limitation an AI agent can have, which is that it may be articulate about everything and able to do nothing.
How do you know the handoff to a human is any good?
An agent that hands off well is worth more than one that hands off rarely. Judge handoffs on three things. Timing: did the agent escalate at the first clear signal of confusion or frustration, or only after several failed attempts wore the customer down? Context: did the human receive the transcript, the customer record, and a short summary, or did the customer have to start the story over? Honesty: did the agent say it was bringing in a person, or did it simply go quiet while the customer waited?
Track requested handoffs separately from automatic ones. A customer who types the equivalent of let me talk to someone is telling you something different from an agent that recognized its own uncertainty and stepped aside. Rising requested handoffs usually point at trust, tone, or a knowledge gap. Rising automatic handoffs on a narrow set of topics are often good news, because the agent is drawing a boundary in the right place. Corebee is built around this behavior, handing off when a customer asks or when it is unsure rather than guessing, and the reason that design choice matters is that it keeps your escalation numbers interpretable instead of noisy.
What should you track channel by channel?
Blended, cross-channel reporting hides more than it reveals. An agent can be dependable on web chat and weak on voice, or excellent on email and clumsy on social messages where customers write in fragments. Measure resolution, escalation, and sentiment per channel, then compare.
Voice deserves its own scrutiny. Callers cannot scroll back, they get one pass at understanding an answer, and they will hang up rather than rephrase. Track abandoned calls, calls that ended without a next step, and how often the caller asked for a human. Disclosure belongs in the same review: an AI receptionist should tell callers it is an AI, both because people notice and because a customer who knows what they are speaking to asks clearer questions.
Messaging channels bring their own pattern. Conversations on WhatsApp, SMS, and Telegram tend to stretch across hours or days with long gaps, so a follow-up window tuned for live chat will mislabel a normal pause as a resolution. Set channel-specific windows, or accept that your messaging numbers are optimistic.
How can you tell whether the knowledge base is the real problem?
Most AI support failures are content failures wearing a technology costume. The model is rarely inventing a wrong policy out of nowhere; more often your policy exists in three versions, two of them outdated, and one of them only in someone's head. The diagnostic is straightforward: collect every conversation where the agent could not answer, hedged, or escalated for lack of information, and group them by topic. That gap list is the most valuable report your setup can produce, because each line is a specific article someone can write this week.
This is also where you find the reverse problem, which is content the agent answers from confidently but that no longer reflects how you operate. Sort your source material by how long it has gone untouched, and review the pieces the agent cites most often. Because Corebee answers from your website and your knowledge base, improving the source material is usually a faster path to better answers than changing any setting, and it improves what your human team sends out too.
How do you compare cost when vendors price so differently?
The metric that survives comparison is cost per resolved conversation, using your own definition of resolved. Getting there means understanding what each pricing model punishes. Per-agent pricing, the familiar shape at Zendesk, Freshdesk, and Help Scout, punishes adding people, so it quietly discourages having a colleague dip into the queue during a busy week. Resolution-based AI pricing, used by Intercom, ties your bill to the thing you are trying to increase, which is coherent but makes budgeting harder as you grow. Gorgias bills around ticket volume and is shaped for ecommerce, which fits stores well and fits professional services awkwardly. Tidio scales by conversation tiers, which suits smaller sites. Chatwoot removes license cost if you self-host and replaces it with engineering time, a genuine trade rather than a free option.
None of these is wrong, but each one changes which behavior you can afford to measure. Corebee is flat $99/month with unlimited seats and unlimited AI conversations, so cost per resolved conversation falls as volume grows rather than tracking it, and nobody has to think about seat count before inviting a teammate to look at a thread. You can see the full terms on the pricing page, and if you are mid-comparison, a structured look at Zendesk alternatives is a reasonable next step.
What does a useful monthly review look like?
Keep it short and repeatable, because a review that takes an afternoon will be skipped by the second month. Open with the anchor number, full resolution rate, and compare it to the previous period. Look at escalations by reason and ask whether the largest reason is a content gap, a missing capability, or a genuine judgment call that should stay with a person.
Then read transcripts. Pick a handful at random, plus every conversation that ended in poor sentiment, plus the longest conversations in the period, since length usually marks confusion. Reading them together, rather than in isolation during the week, is what turns scattered irritations into a pattern you can name.
Close by writing down one change and one thing you expect it to move. If you rewrite the returns article, say that you expect returns escalations to fall. Next month you will either be right, which builds confidence in your measurement, or wrong, which teaches you something about where the problem actually lives. Over a few cycles this habit does more for support quality than any change of vendor, because it forces your definition of working to stay in contact with what customers experience.