Why does an AI support tool invent answers at all?
A language model is built to produce the most plausible next sentence, not to admit a gap. Left unconstrained, a question about your refund window gets answered with what refund windows usually look like, written in the same confident tone as a correct answer. That is the whole problem, and it is a design choice rather than something you can train away on your own.
In a support setting it usually traces back to one of three causes. Retrieval is weak, so the tool searches your content, finds nothing that matches closely, and answers from general knowledge instead of stopping. Or the instructions are loose: the assistant is told to use your articles as helpful context rather than being restricted to them. Or your documentation never covered the case at all, so there was nothing true to retrieve and the model filled the space.
A fourth case gets blamed on invention but is something different. If an article is out of date, a properly grounded assistant repeats it faithfully and the customer still receives a wrong answer. Grounding addresses fabrication. Stale content is yours to maintain, and no vendor can clean it up on your behalf.
What does grounding in a knowledge base actually mean?
Vendors use the word loosely, so it helps to break it into four behaviors you can check during a trial.
Retrieval before generation. The tool searches your help center, website and uploaded documents, pulls the passages that match the question, and writes only from those passages. If retrieval comes back with nothing strong, generation should not proceed at all.
Attribution. The answer points back at the article it came from. For your team this is the fastest way to audit a bad reply, because you can tell immediately whether the assistant misread the article or the article itself was wrong. Those two problems have completely different fixes.
Abstention. There is a point below which the assistant declines rather than guesses. This is the behavior most often missing, and it is the one that separates a tool that sounds careful from a tool that is careful.
Escalation. Declining only helps if something happens next: a handoff to a person, a ticket with the full conversation attached, a callback booked. An assistant that says it cannot help and then goes quiet has moved the failure rather than removed it.
One more control is worth asking about: scope rules. You should be able to tell the assistant which subjects it must never address, such as discounts, legal questions, medical advice or anything that commits your business to a promise, regardless of what sits in the knowledge base.
Which AI support tools keep answers tied to your own content?
Every serious vendor now retrieves from your content. The useful differences are in controls, channel coverage and the shape of the price.
Zendesk's AI agents answer from help center articles inside a mature admin and reporting stack, which suits teams that already run Zendesk and need granular permissions and audit trails. Intercom's Fin draws on help center articles plus additional content sources you point it at, and shows its sources; it bills by resolution, so cost tracks volume. Freshdesk pairs Freddy with its solution article structure and fits teams already standardized on the wider Freshworks suite. Gorgias is built around ecommerce and can combine policy pages with order data, which matters when the correct answer depends on the specific order in front of it. Chatwoot is open source and can be self-hosted, so teams with technical staff keep data and model choices in their own hands. Help Scout takes a docs-first, deliberately simple approach that works for small teams with no dedicated admin. Tidio's Lyro targets small businesses and works from site content and question-and-answer pairs.
Corebee sits in the same grounded category: it answers from your website and knowledge base, hands off to a human when the customer asks or when it is not confident, and applies the same behavior across web chat, email, voice, WhatsApp, Slack, Facebook, Instagram, SMS and Telegram, with setup that takes about ten minutes and no developer. The retrieval side is described on the knowledge base page, and handoffs land in the shared inbox.
How can you test a tool for invented answers before you buy?
Do not evaluate on questions your documentation answers plainly. Every vendor passes those. Build a list with four kinds of questions and run it identically across each tool you are considering.
Baseline questions, which your articles cover clearly. Uncovered questions, which you know are absent from your content, because this is where fabrication shows up. Near-miss questions, worded closely enough to an existing article that retrieval will match it while the actual question is different, such as asking about returns on a sale item when your article only covers full-price returns. Pressure questions, where the customer asks for an exception, a discount, a delivery promise or a comparison with a competitor.
Score each reply on four points: did it answer from a real article, did it show the source, did it decline cleanly when it had nothing, and did it route the conversation to a person. Then repeat the list on every channel you plan to use. Voice deserves the harshest look, because the caller sees no citation and cannot scroll back to check. Run the list again after you edit your content, since a rewritten article changes what retrieval finds.
What does your knowledge base need to look like for this to work?
Grounding is only as good as the material it is grounded in, and most hallucination complaints in small teams turn out to be content problems wearing a different costume.
Keep one fact in one place. When the same policy appears in three articles with slightly different wording, retrieval picks one and you cannot predict which. Write down the edge cases rather than leaving them in a manager's head: who can approve an exception, what the cutoff times are, what happens on holidays. Use the words customers use in headings, not internal product names. Separate internal notes from public content, so pricing logic and escalation rules never surface in a customer-facing answer.
Then close the loop. Every conversation the assistant declined is a gap list handed to you for free, and the fastest improvement cycle is turning those declines into short articles each week.
What should happen when the assistant is not sure?
Decide the unsure path per channel before launch, because the right move differs. In chat, the assistant should hand over to a person in the same thread or open a ticket with the transcript attached. In email, an uncertain answer can sit as a draft for a human to approve, which costs you a little time and removes the risk entirely. On voice, the agent should identify itself as an AI, take the question, and either transfer or book a callback rather than improvising. On WhatsApp and other messaging channels, where customers expect asynchronous replies, a clear promise of a human follow-up is usually enough.
One rule applies everywhere: if a customer asks for a person, that request should be honored on the first ask, with no deflection loop in between. An assistant that argues about it does more damage to trust than an occasional wrong answer.
Does pricing change how safely you can run an AI agent?
Indirectly, yes, and it is worth thinking through before you sign. Per-resolution pricing ties your bill to deflection, which means you need to understand what the vendor counts as a resolution, and it can create quiet pressure against the conservative behavior you actually want, since a careful decline that becomes a human conversation may be charged on both sides. Per-seat pricing makes adding reviewers expensive, so teams keep fewer people in the loop than the work deserves, precisely when a second pair of eyes is cheapest insurance.
Flat pricing removes both tensions. Corebee is $99 per month, flat, with unlimited seats and unlimited AI conversations, no per-seat and no per-resolution fees, so handing a conversation to a person is a judgment call rather than a budget decision. You can compare the structures on the pricing page and in the Zendesk alternatives breakdown.