What does losing context actually look like?
When a handoff breaks, the customer notices before you do. They have already given their order number, explained the problem twice, and been told someone will take over. Then a person arrives and opens with "Hi, how can I help?" and the customer starts again from the beginning. Sometimes it is worse than repetition: the bot quoted a policy or promised a callback, the agent has no idea that happened, and the business now appears to contradict itself in the same window.
There is usually an architectural reason behind it. Many AI chat tools are attached to a help desk from the outside. The bot lives in one vendor's widget, the humans work in another vendor's ticket queue, and the bridge between them is an integration that opens a new ticket containing a short summary. A summary is a lossy copy. It loses the exact wording, the order the questions came in, the links the bot sent, and the tone clues an experienced agent reads instinctively. It often loses the thread as well, so the agent's reply leaves as a fresh email while the customer is still watching the chat box.
The second common cause is identity. If the bot spoke to an anonymous visitor and the help desk organizes everything around an email address, the conversation has nothing to attach itself to. The agent sees words with no relationship behind them: no past orders, no open ticket, no note that this is the third time this week. That is context loss too, just a quieter kind.
What has to travel with the conversation for a handoff to hold?
It helps to be specific about what "context" means, because vendors use the word loosely. A handoff holds together when all of the following move with it.
The verbatim transcript, not a summary. The agent should be able to scroll up and read what the customer actually typed, including the failed attempts. A generated summary is useful as a header, never as a replacement.
Identity and history. Email, phone number or social handle, plus whatever the bot already collected, linked to a single customer record so past conversations are one click away.
The same thread on the same channel. If the conversation started on WhatsApp, the agent's reply has to arrive on WhatsApp, in the same thread, under the same business number. Replying through a different channel is a new conversation as far as the customer is concerned.
What the bot asserted. Any policy it quoted, article it linked, price it stated or commitment it made. This is the part most integrations drop, and it is the part that creates awkward conversations.
Structured data the bot captured. Order number, account ID, booking reference, the reason for contact. If the bot asked for it, no human should ask again.
Why it escalated. "Customer asked for a person," "repeated the same question after two answers," and "topic outside the knowledge base" call for different opening lines from the agent.
The customer's wait state. How long they have been waiting, and whether they were told someone would arrive now or reply later.
A product either carries these by design because the bot and the inbox are one system, or it reassembles them through an API and drops some in transit. You can tell which you are looking at within a few minutes of testing.
How do the main AI chatbot tools handle this today?
Fair comparison matters more than a scoreboard here, because these tools were built for different buyers.
Intercom built its AI agent directly on its own inbox, so escalation is a thread handed to a teammate rather than a ticket copied from elsewhere. Context retention is strong. The trade-off for smaller teams is commercial: seats are priced per user and AI work is often billed per resolution, so both busy months and growing teams raise the bill.
Zendesk offers a mature agent workspace, deep ticketing, routing rules and reporting, and its AI agents escalate into that workspace with the conversation attached. It suits teams that need formal queues, SLAs and audit trails. It also carries the most configuration overhead of the group, and its AI capability commonly sits on higher tiers alongside per-seat costs. If that structure is the reason you are shopping, a side-by-side look at Zendesk alternatives is a reasonable next step.
Freshdesk with Freddy follows a similar shape at a gentler learning curve, with solid ticketing fundamentals and a familiar agent view.
Gorgias is built for ecommerce, particularly Shopify, and its advantage is commercial context: order, refund and subscription data sits beside the message, so an escalated chat arrives with the things a store agent actually needs. Outside retail, that focus is less of an advantage.
Help Scout keeps the shared inbox and help center simple and pleasant, and handoff within it is clean because there is not much machinery to cross. Teams wanting heavy automation may outgrow it.
Chatwoot is open source and self-hostable, which appeals if you want to own the data and the deployment. The context question becomes partly your own engineering question, since much depends on how you wire the AI layer to it.
Tidio targets small stores with live chat plus its Lyro assistant, and handoff into the same chat console is straightforward. Larger support operations tend to hit limits on routing and reporting.
Across the category, two patterns hold. Tools where the AI and the human inbox are the same product keep context; tools stitched together by integration lose some of it. And handoff quality is a pricing question as much as a technical one, because per-seat and per-resolution models quietly discourage the escalations that keep customers happy.
What should trigger a handoff, and who decides?
A bot that never escalates is worse than one that escalates too often. Set explicit triggers rather than trusting a confidence score alone.
Start with the direct request: any phrasing close to "talk to a human" should route immediately, with no attempt to answer one more time. Add topic triggers for situations no bot should own, such as cancellations, chargebacks, complaints about being charged twice, legal threats, safety issues and anything involving a distressed customer. Add a repetition trigger, because a customer rephrasing the same question twice has already told you the answer is not landing. Add an uncertainty rule so that when the answer is not clearly supported by your knowledge base, the bot says so and passes the conversation on instead of improvising.
Then decide what happens when nobody is there. Out of hours, the honest move is to say so, take the details, and put the conversation in a queue that a person actually opens in the morning. The worst pattern is a promise of immediate help followed by silence, and it is common in setups where the escalation creates a ticket in a system the small team rarely checks.
Does context survive when the channel changes?
This is where most products separate. Web chat handoff is the easy case. The harder cases are the ones customers actually use.
On WhatsApp, the conversation is a persistent thread tied to a phone number, so a handoff that arrives as an email breaks it visibly. The agent needs to reply inside that thread from the same business number, which means the inbox has to hold the channel natively rather than forward a notification. The same applies to Instagram and Facebook messages, SMS, Telegram and Slack, each with its own threading rules. A product that treats WhatsApp as a first-class channel will keep the thread; one that treats it as a webhook usually will not.
Voice is the hardest of all, because the context has to be created before it can be passed. A phone call produces no transcript unless something transcribes it, and a transfer can drop the caller into a queue with nothing attached. Corebee's voice channel is an AI receptionist that tells callers it is an AI, answers from the same knowledge the chat channels use, and passes the call to a person when asked or when it is unsure, with the conversation kept in the shared inbox rather than in a separate call log. Whatever tool you choose, ask two questions about voice: does the caller know they are speaking to an AI, and does the human who picks up see what was already said.
Cross-channel continuity deserves a check of its own. A customer who chatted on the website yesterday and calls today should not have to re-explain, which only works if both conversations attach to the same record in the same inbox.
How do you test a handoff before your customers do?
Do not evaluate this from a feature table. Run it.
Take a handful of conversations your team genuinely receives, including one angry one and one where the answer is not in your documentation. Run each of them through the bot as a customer would, on every channel you plan to support, using your phone rather than a desktop preview. At the moment of handoff, note exactly what the agent sees: full transcript or summary, customer identity present or missing, escalation reason visible or absent, and the channel correctly identified.
Then reply as the agent and confirm the customer receives it in the original thread. Check whether the conversation appears against the right contact in your CRM, whether an internal note survived, and what a second agent sees if the first hands it on. Finally, repeat the test after hours, which is when the gaps usually appear.
One practical detail gets missed: seat access. If the person who should take over does not have a login because seats cost money, the handoff fails for commercial reasons rather than technical ones. Corebee is flat $99 per month with unlimited seats and unlimited AI conversations, and no per-resolution fee, which exists precisely so that adding the warehouse manager or the weekend cover person to the inbox is not a budget decision. You can see the full breakdown on the pricing page.
What should a small team prioritize when choosing?
Rank the criteria in this order. First, is the AI and the human inbox one system or two? Everything else follows from that answer. Second, are your real channels supported natively, including voice and WhatsApp if customers use them, with replies going back into the same thread. Third, does escalation happen on explicit request without argument. Fourth, can you configure what the bot refuses to handle, rather than hoping the model decides well. Fifth, does the pricing model penalize either escalation or extra teammates, since both are things you will want more of as you grow.
Setup effort is worth weighing too. A tool your support lead can configure without a developer, connect to your website content and CRM, and adjust after the first week of real conversations will outperform a more capable platform that sits half configured because nobody had time to finish it. The best handoff is the one your team will actually maintain.