Why support costs actually grow
Key Takeaways
For busy support leads: before you write a hiring case, write this formula on a whiteboard: customers times conversations per customer times minutes per conversation times cost per minute. Then find your real number for each. Most teams discover their minutes per conversation is far higher than they assumed, because nobody counts context switching, follow-ups or note taking. That single measurement usually changes the decision.
- 1Deflection is the cheapest lever. A conversation that never starts costs nothing to handle and nothing to automate.
- 2Automation is the biggest lever. It removes whole categories of human work rather than shaving minutes off each one.
- 3Process work has a floor. You cannot cut below the time it actually takes to think about a hard problem.
- 4Headcount is a capability lever, not a volume lever. Hire for a new language, channel or specialism, not for a queue.
- 5Capacity is not cash. Savings only appear on the P&L if you genuinely do not make the next hire.
Almost every scaling article starts from the customer-to-agent ratio. It is the wrong starting point, because the ratio is an output, not an input. You cannot manage it directly, and comparing yours to somebody else's tells you nothing about products with different complexity.
Model the inputs instead:
Monthly support cost = customers x conversations per customer x minutes per conversation x loaded cost per minute
| Multiplier | Who owns it | How you move it |
|---|---|---|
| Customers | Sales and marketing | You do not want to move this one down |
| Conversations per customer | Product, docs, proactive comms | Fix the causes: confusing UI, unhelpful error messages, missing self-service |
| Minutes per conversation | Support ops and tooling | Automate whole conversations, then cut the minutes around the ones that remain |
| Loaded cost per minute | Finance and staffing model | Tiering and location strategy, with real limits and real trade-offs |
Growth is only a cost problem if you treat the middle two as fixed. They are the least fixed things in the whole business.
A worked example: what happens when customers triple
Every number below is an illustrative assumption, not a benchmark. Substitute yours.
Start with 2,000 customers who each open 0.5 conversations a month, so 1,000 conversations. Assume 9 minutes of real work per conversation, counting the reply, the follow-up, the context switch and the notes. Assume a loaded cost of $35 per agent hour. Assume one full-time agent delivers about 120 productive support hours a month once you subtract meetings, training and everything else.
That baseline is 150 hours, $5,250 a month, and 1.25 agents' worth of work.
Now triple to 6,000 customers.
| Step | Conversations | Minutes each | Agent hours | Labour at $35/hr |
|---|---|---|---|---|
| Today, 2,000 customers | 1,000 | 9 | 150 | $5,250 |
| 3x customers, nothing else changes | 3,000 | 9 | 450 | $15,750 |
| Lever 1: conversations per customer 0.5 to 0.35 | 2,100 | 9 | 315 | $11,025 |
| Lever 2: AI closes half the rest | 1,050 | 10 | 175 | $6,125 |
| Lever 3: 10 minutes down to 8 | 1,050 | 8 | 140 | $4,900 |
| Add quality review and content upkeep | 12 | $420 | ||
| After all three levers | 152 | $5,320 |
Add a flat $99 tool and the total is $5,419.
Read those two outcomes side by side. Doing nothing costs $10,500 a month more than today. Pulling all three levers costs about $170 a month more than today, for three times the customers. Headcount goes from roughly 1.25 to 1.27 rather than to 3.75.
Two honest notes. Minutes per escalated conversation went up to 10 in the Lever 2 row, because an agent now reads what the AI already tried. And the levers do not all land at once. Deflection compounds over quarters, automation lands in weeks, process work is continuous.
Lever 1: stop the conversation from starting
The highest-return work in support is usually not in support. Pull your top five conversation drivers each month and ask, for each one, why this is a conversation at all. Nearly always the answer is a product decision: an ambiguous label, an error message that states a failure without stating a fix, a billing event with no explanation attached, or a feature people cannot find.
The order of return, in our experience running this loop:
- Fix the product or the error message. The conversation stops permanently and costs nothing forever after.
- Answer it in the interface, at the moment of confusion, rather than in an article somebody has to go and find.
- Send it before they ask. Status pages, incident notices, plan change confirmations and expiry warnings kill entire spikes.
- Write the article. Necessary, but the weakest of the four, because it depends on the customer searching.
Most teams do these in exactly the reverse order, because writing an article is the only one support can do without asking another team for anything.
Lever 2: resolve documented questions automatically
Once a conversation exists, the cheapest resolution is one with no human in it. This is what AI support is genuinely good at: questions with a documented answer, asked in a customer's own words, at 3am on a Sunday.
Be precise about what it removes. It removes whole conversations from the queue in the categories you allow it to handle, and it shaves drafting time off the ones it does not. It does not remove judgment, investigation, or the responsibility to answer angry people quickly with a human.
The measured effect on human productivity is more modest than the case studies imply. In a study of 5,179 support agents at one software firm, access to an AI assistant raised issues resolved per hour by 14 percent on average, with a 34 percent improvement for novice and low-skilled workers and minimal impact on experienced ones (Brynjolfsson, Li and Raymond, NBER, 2023). Compare that with the number vendors circulate: Klarna reported that in its first month its AI assistant handled "2.3 million conversations, two-thirds of Klarna's customer service chats", doing "the equivalent work of 700 full-time agents". That is a self-reported, unaudited, first-month figure the company itself walked back in 2025. Build your model on the peer-reviewed one.
It also adds work, which is the part the vendor arithmetic omits. Somebody samples answers for quality. Somebody keeps the content current. Agents spend an extra minute reading the AI transcript on every escalation. In the worked example above that overhead was 12 hours a month, and if you do not budget for it, it comes out of the same team you were trying to relieve.
Lever 3: cut the minutes around every human reply
For the conversations that still need a person, the target is not the thinking time. It is everything wrapped around it: searching for the answer, switching tools, re-typing the same explanation, deciding who should own it, and updating records afterwards.
What reliably works: templates for your top 20 response patterns, kept as starting points rather than scripts. An internal knowledge base separate from the customer-facing one, holding escalation procedures and known edge cases. Routing that matches specialism to topic, so a billing question does not land with someone who has to go and learn billing. Batching similar conversations rather than alternating between billing, technical and onboarding.
This lever has a floor, and pretending otherwise is how teams burn people out. Once you have removed the waste, the remaining minutes are the work. Pushing further just means worse answers delivered faster.
Where the savings quietly leak back out
| Leak | How it shows up | Fix |
|---|---|---|
| The escalation tax | Every AI-touched escalation costs an extra minute of transcript reading | Budget it in the model, and make handoff summaries short and structured |
| Success-priced tooling | Per-resolution billing means the bill grows exactly as automation works | Model your bill at 2x current volume before signing |
| Per-seat sprawl | Adding ops, engineering or product to the helpdesk gets priced out | Prefer flat pricing if you want more people in the tool |
| Review debt | Nobody samples answers after month two, quality drifts silently | Fixed weekly slot, owned by a named person |
| Capacity mistaken for cash | Hours are freed, the hire happens anyway, nothing lands on the P&L | Decide up front what the freed hours are for |
Put real numbers on the second row. As fetched on 2 August 2026, Zendesk's own explainer states it "charges $1.50 per automated resolution", Intercom prices Fin at $0.99 per outcome, and Gorgias charges "$0.90 on most plans" per resolved conversation. At the 1,050 automated conversations in the worked example above, that is roughly $945 to $1,575 a month on the meter alone, and it grows every time your content improves.
The capacity row is the one that turns a good project into a disappointing board slide. Freed hours are only savings if a specific planned cost disappears. Otherwise they are quality, speed and sanity, which are worth having but do not belong in a budget line.
Which lever should you pull first?
| Your situation | Pull first | Why |
|---|---|---|
| Little or no documentation | Lever 1, content | Automation on top of nothing produces confident wrong answers |
| Good docs, high repetitive volume | Lever 2, automation | The highest-return single change available to you |
| Low volume, high cost per conversation | Lever 3, process and tiering | Not enough volume for automation to repay setup |
| Spiky seasonal volume | Lever 1, then flexible staffing | Peaks are a capacity problem, not a permanent headcount problem |
| Mostly bespoke technical investigation | Lever 3, internal tooling | The work is genuinely hard, so target the minutes around it |
The sequencing matters more than the tooling choice. Automating a broken process scales the breakage.
When hiring actually is the right answer
Scaling without hiring is not the same as never hiring. Add people when the work changes in kind rather than in quantity.
A new language or time zone is a hire. A specialism your team does not have, such as deep technical implementation support, is a hire. A backlog that persists after you have pulled all three levers is a hire, because the queue is now genuinely made of human work. Sustained burnout signals, rising sick days, falling internal quality scores and turnover, are a hire, and treating them as a motivation problem is how you lose the people who know your product best.
What is not a hire: raw volume in categories that are documented, repetitive and automatable. Hire for capability, not for queue length.
When cost cutting damages the business
There is one thing you should never optimise, and that is the exit to a human. Hiding it, gating it behind clarifying questions, or making customers argue with a bot to earn it, all reliably save money this quarter and cost customers next quarter. Our position on human handoff is that it should be visible in every conversation and carry the full transcript across, so nobody repeats themselves.
The same applies to answer quality. A cheap system that is confidently wrong generates a second conversation, a trust problem and sometimes a refund. Measure repeat contacts on the same topic within 72 hours alongside every efficiency metric you track. If that number rises while cost per conversation falls, you have not saved anything. You have moved the cost somewhere you are not looking.
The five minute version
Write down the four multipliers, measure your real minutes per conversation, and stop benchmarking your customer-to-agent ratio against companies with different products. Fix the causes of your top five conversation drivers, automate the documented questions, cut the waste around the human replies, and budget honestly for the review and maintenance that automation creates. Then hire deliberately, for capabilities you do not have rather than for a queue you have not yet optimised.
Want to model this against your own volume? Start a free trial and measure the first lever before you commit to any of them.