Why most support benchmarks you read are unusable
Key Takeaways
For busy support leads: if you only change one thing this quarter, stop reporting first response time and start reporting time to first useful answer at the median and the 90th percentile. Auto-acknowledgements make the standard metric look excellent while your customers wait exactly as long as they did before.
- 1Time to first useful answer. Measured from the customer's first message to the first reply that actually moves the issue forward, reported as median and p90.
- 2Reopen rate. The observed version of first contact resolution, because the agent who closed the ticket is not a neutral judge of whether it was solved.
- 3Contact rate per 100 active accounts. Absolute ticket volume is meaningless while your customer count is moving.
- 4Cost per resolved conversation. Useful only when you state the denominator out loud, because the choice of denominator changes the number by more than half.
- 5Escalation quality. The rate tells you nothing without knowing which direction the mistakes run.
An earlier version of this article carried benchmark ranges: CSAT of 85 to 92 percent for B2B SaaS, first contact resolution of 70 to 80 percent, cost per resolution of 8 to 15 dollars. Each was attributed to a research landing page. We removed all of them, because we could not trace a single figure to a primary study with a stated method. If we could not verify our own numbers, you certainly could not.
Here is how these numbers get laundered. A vendor publishes a survey. A blog cites the vendor but rounds the figure and drops the sample description. A second blog cites the first blog. By the fourth hop the number has a decimal point, a year, and no methodology. Search for any support benchmark and you will find the same range repeated with four different attributions.
Even where the underlying research is real, definitions destroy comparability. Does first response time include your autoresponder? Does a conversation abandoned by the customer count as resolved? Is CSAT measured on a five point scale with the top two boxes counted, or a thumbs up? Two teams reporting 82 percent can be measuring genuinely different things.
If you want an outside reference point, use one that publishes its method. The American Customer Satisfaction Index runs roughly 200,000 interviews a year across more than 400 companies, and reported that in the first quarter of 2026 the index fell 0.3 percent to a score of 76.7, with customer complaints surging by 16 percent (ACSI, 2026). That tells you which way the national weather is blowing. It still does not tell you what your team should hit.
The research worth quoting is the research about mechanisms rather than levels. The Corporate Executive Board work published in Harvard Business Review found that 96 percent of customers who had a high effort service interaction became more disloyal, against only 9 percent of those with a low effort experience (Harvard Business Review, 2010). That finding tells you what to optimise. It does not tell you what number your team should hit, and no honest source will.
Metric 1: time to first useful answer
Define it precisely: the elapsed time from the customer's first message to the first reply that advances the issue. An automated acknowledgement does not count. Neither does "thanks, we are looking into this." If the reply does not answer the question, ask a question that unblocks the answer, or state a specific next step with a time, the clock is still running.
Standard first response time is the easiest metric in support to fake. Turn on an autoresponder and it collapses to under a minute overnight while the customer experience is completely unchanged. We have watched teams celebrate that chart.
Report median and 90th percentile, never the mean. One ticket that sat over a holiday weekend will drag a mean into uselessness. A median of 6 minutes with a p90 of 9 hours is not a speed problem, it is a coverage problem, and the fix is a rota or an out of hours answer, not more urgency during the day.
Metric 2: reopen rate, not first contact resolution
First contact resolution is usually recorded by the agent who closed the conversation, which makes it a self-graded exam. Reopen rate is observed behaviour: the percentage of conversations that receive a new customer message within seven days of being marked resolved, counting new threads about the same topic as reopens.
Work the arithmetic. A team handling 900 conversations in a month with 74 reopens has a reopen rate of 8.2 percent. That single number is mildly interesting. The segmentation is where the value sits: if 22 of those 74 reopens are billing conversations and billing is only 9 percent of total volume, billing is generating reopens at roughly three times its share. You now have one project instead of a vague instruction to improve quality.
Do not attach a cost multiplier to reopens unless you have measured yours. The common claim that a reopen costs double has no source we can find. What is safe to say is that a reopen costs a second full handle plus the time an agent spends reloading context, and it arrives with a customer who is already annoyed.
Metric 3: contact rate per 100 active accounts
Total ticket volume is only readable if nothing else about your business is moving, which is never. Contact rate per 100 active accounts separates growth from deterioration.
Take the same team. Month one: 900 conversations across 420 active accounts, which is 2.14 contacts per account. Month four: accounts have grown 20 percent to 504 while conversations grew 32 percent to 1,188. Contact rate is now 2.36, up 10 percent. Raw volume said "we grew." Contact rate said "each customer needs us more than they used to," which is a product or onboarding signal that raw volume hid completely.
One false alarm to watch for. If you change the definition of an active account, for example by counting invited-but-never-logged-in seats, the denominator moves and the metric lies for a quarter. Freeze the definition in writing, with a date, before you start trending it.
Metric 4: cost per resolved conversation, and the denominator argument
This is the metric that connects support to the P&L, and it is almost always quoted without the one detail that determines the answer.
Work it through. Three support people at a fully loaded cost of 58,000 per year is 174,000 per year, or 14,500 per month. Add 600 per month of tooling and you have 15,100 in monthly support cost. The team handled 900 conversations, of which 360 were resolved by AI or self-service with no human involvement, leaving 540 human-touched.
| Denominator | Calculation | Result | What it answers |
|---|---|---|---|
| All conversations | 15,100 / 900 | 16.78 | What a conversation costs the business today |
| Human-touched only | 15,100 / 540 | 27.96 | What one more human conversation costs at the margin |
Same month, same team, two numbers that differ by 67 percent. Use the human-touched figure for staffing and automation decisions, because that is the cost you actually avoid when a conversation is deflected. Use the all-conversations figure when you are comparing your total support cost to revenue. Never quote either without saying which one it is.
There is a pricing consequence people miss. If your support platform charges per resolution, your cost per resolved conversation has a floor that rises with your own success, and every deflection you engineer partially benefits your vendor. Flat pricing does not behave that way. Corebee charges a flat 99 dollars per month for exactly this reason, and you can see the full structure on our pricing page.
Metric 5: escalation quality, not escalation rate
An escalation rate on its own is unreadable. Twenty percent could mean your AI and your first line are handling the right things, or it could mean they are guessing at questions that should have gone to a specialist an hour ago.
There are two error types and they are not symmetrical:
Over-escalation is wasteful. A conversation gets handed to a human that the knowledge base already answered. You lose an agent's time and the customer waits longer than necessary.
Under-escalation is dangerous. A refund dispute, a security question, or an angry churn-risk account gets a confident automated answer instead of a person. One of these costs more than a hundred over-escalations.
Measure it by sampling, not by counting. Every week, pull 20 escalated conversations and 20 that were resolved without escalation, and label each one as correctly handled or not. That is roughly 40 minutes of a lead's time and it produces the only escalation number worth acting on. Then fix the rules rather than the people. Our note on designing the human handoff covers the trigger conditions worth hard-coding.
The six metrics we stopped reporting
| Metric | What it rewards | What it costs you |
|---|---|---|
| Average handle time | Ending conversations quickly | Rushed answers that come back as reopens |
| Tickets closed per agent | Volume over correctness | Premature closes and cherry-picked easy tickets |
| CSAT as a headline number | Whatever the small self-selecting minority felt | False calm, since quietly disappointed customers rarely respond |
| Support NPS | A number nobody can act on | A quarterly debate with no owning team |
| Total ticket volume alone | Looking good while shrinking | Blindness to rising contact rate during growth |
| Vendor-reported auto-resolution rate | The vendor's definition of resolved | Abandoned conversations counted as wins |
None of these are worthless. Average handle time is genuinely useful for capacity planning, and CSAT is a reasonable smoke alarm when it moves sharply. The point is that none of them belong on the dashboard the team looks at every week, because each one rewards a behaviour you do not want.
The NPS row has research behind it rather than just our opinion. A longitudinal study of 21 firms and more than 15,500 interviews, published in the Journal of Marketing, failed to replicate the claim that Net Promoter is a superior predictor of growth (Journal of Marketing, 2007). The score is not worthless. It is simply not the special instrument it is marketed as, and a support-level version of it is one further step removed from anything a team can act on.
The auto-resolution row deserves an extra word, given what we sell. If a customer asks a question, the bot answers, and the customer leaves without replying, many systems count that as resolved. It might have been resolved. It might have been surrender. Until you can separate those two by checking whether the same customer opened a new conversation about the same topic within a week, treat the headline number as marketing.
Where this framework breaks
Below roughly 200 conversations a month, percentages are noise. At 120 conversations, four reopens instead of no reopens moves your reopen rate by 3.3 points, which will look like a trend on a chart and is nothing at all. Small teams should read conversations, not rates. Twenty full transcripts a week will teach you more than any dashboard until volume gets serious.
Contact rate per account misleads badly when account sizes vary wildly. A 4,000 seat customer and a 3 seat customer are not one unit each. If that is your shape, use contacts per 100 seats, or per active workspace, and segment by tier.
Seasonality wrecks month-over-month comparison in any business with a fiscal or academic cycle. Compare to the same period last year once you have the history, and until then annotate your charts with what happened, because in six months nobody will remember the migration that caused the March spike.
How often to look at each number
| Cadence | What you look at | The decision it triggers |
|---|---|---|
| Daily | Oldest unanswered conversation, queue size | Move someone, or accept the wait |
| Weekly | Time to first useful answer p90, reopen rate by topic, 40 sampled escalations | Pick one topic to fix this week |
| Monthly | Contact rate per 100 accounts, cost per resolved conversation | Staffing, automation scope, pricing conversations |
| Quarterly | Definition audit: has anything changed under the metric | Re-baseline and say so publicly |
The quarterly definition audit is the step everyone skips. Metrics rot silently when someone changes a status name or adds a channel.
How to build your own baseline in two weeks
Week one: write the definitions down before you pull any data. What counts as a conversation, when the first useful answer clock stops, what a reopen is, which costs go into the cost line. Half a page. Circulate it and let people argue, because the argument is where you discover that two people on your team have been counting differently for a year.
Week two: pull four weeks of history against those definitions, by hand if necessary. Then publish the five numbers with no commentary and no targets. Targets in week two are guesses. After eight weeks of your own data you will know what normal looks like for your product, your customers and your team, which is the only benchmark that can genuinely tell you whether you improved.
Then pick one number and move it. Not five. One. The teams that get value from measurement are the ones that treat a dashboard as a queue of experiments rather than a report card.
If you want to see these numbers without building the plumbing yourself, try Corebee free.