What does it mean for an AI to ground answers in your documentation?
Grounding is a retrieval step that happens before the model writes a single word. The tool takes your help articles, policy pages, product descriptions and whatever else you point it at, splits them into passages, and stores them so they can be looked up by meaning rather than by exact keyword match. When a customer asks something, the system searches that store, pulls the passages that look most relevant, and hands them to the language model with an instruction along the lines of "answer using only this material."
That instruction is the whole point. A model answering from general training knowledge will produce something fluent about return windows in general. A grounded model produces something about your return window, because your return policy is sitting in front of it. The practical test is whether the answer changes when you edit the source article. If you rewrite the article and the assistant keeps repeating the old line, then whatever the vendor calls the feature, it is not reading your documentation at the moment it answers.
Two capabilities normally travel with grounding, and you should treat them as part of the same question. The first is citation: the reply links to the article it drew from, so the customer can read further and your support lead can audit the reply afterwards. The second is refusal: when retrieval comes back with nothing useful, the tool says it does not know and offers a human instead of improvising. A tool with retrieval but no refusal behavior will still invent things, just with a more convincing accent.
Which tools ground answers in your own content, and how does each go about it?
Almost every established help desk now has a retrieval layer, so the useful comparison is about the shape of source material each one expects and how much rework that implies for you.
- Zendesk builds on its help center. Its AI agents answer from published articles, which suits teams that already maintain a structured knowledge base inside Zendesk and want answers, macros and ticket routing in one place. If your documentation lives elsewhere, expect a migration or a sync to be part of the project.
- Intercom takes a broader view of sources. Fin can draw on help center articles, public web pages, uploaded files and internal snippets written specifically for the AI, which helps when your real answers are scattered across places that were never tidy articles.
- Freshdesk grounds Freddy in its solution articles, with a similar trade-off to Zendesk: strong if you are already inside the Freshworks suite, more setup if you are not.
- Gorgias is built for ecommerce and its grounding reflects that. Alongside help content it leans on store data such as order status, so its answers are often about one specific order rather than a general policy.
- Help Scout pairs its AI features with Docs, its own knowledge base product, and aims at a plain, conversational tone. Teams that enjoy writing clear articles tend to get good results from it.
- Tidio grounds Lyro in a knowledge base or FAQ set that you supply, and is a common choice for smaller ecommerce teams focused on on-site chat.
- Chatwoot is open source, which changes the calculus. You can ground its assistant in your own articles and, if you self-host, keep the pipeline on infrastructure you control. The cost is that more of the setup and upkeep belongs to you.
- Corebee reads your website and your knowledge base rather than requiring you to rebuild your content inside a new help center first, which is why setup usually takes about ten minutes and no developer. It hands off to a human when a customer asks or when it is not confident, and it can book meetings and write the conversation back to your CRM.
None of these descriptions should be taken as the last word. Retrieval features change often, so check current vendor documentation for the specific sources and limits that apply to the plan you are considering.
Where does the source content come from, and how does it stay current?
There are broadly two ingestion styles, and most tools do some of both. One is crawling: you give the tool a domain and it reads what is published. This is fast to start and good if your website already carries your real answers, including shipping terms, pricing pages and product detail. The other is curated articles: you write or import a set of help articles and the tool answers only from those. Curation gives tighter control over wording and reduces the chance of an old marketing page contradicting your current policy.
The question to press a vendor on is refresh. When you change an article, does the answer change on the next conversation, or after a scheduled re-index? For a small team this matters more than it sounds, because your policies change quietly. Someone updates a shipping cutoff in a page footer and nobody tells support. If the AI reads the live source often, that edit propagates. If it reads a snapshot taken when you onboarded, you have built a machine that confidently repeats last quarter's rules.
The other half of currency is your own hygiene. Contradictory documentation is the most common cause of bad grounded answers, and no vendor can resolve it for you. If two pages describe the same policy differently, retrieval will sometimes find one and sometimes the other. Give each policy one canonical home, delete or redirect the duplicates, and name one person who owns the knowledge base the way someone owns the pricing page.
Why do grounded answers still go wrong?
Grounding narrows the failure modes rather than removing them. The ones worth planning for:
Gaps. The question is reasonable and the documentation simply does not cover it. A well-behaved tool escalates. A poorly configured one stretches a nearby passage into an answer that sounds official and is not.
Vocabulary mismatch. Your article says "subscription cancellation." Your customer types "how do I stop getting charged." Semantic search usually bridges that, but not always, especially with industry jargon or product nicknames that appear nowhere in your writing.
Compound questions. "Can I change the address on an order that already shipped, and will that reset the warranty?" needs two passages combined correctly. Partial retrieval produces a reply that is half right, which is worse than a handoff.
Account-specific questions dressed as policy questions. "Where is my order" cannot be answered from documentation at all. It needs a data lookup. Tools that treat every question as a retrieval problem will answer the general policy and leave the customer no better off.
Confidence without grounding. Watch for replies with no citation. If the tool can produce an answer that no passage supports, it can produce a wrong one.
How do you test grounding before you commit?
Do not evaluate on invented questions. Pull real conversations out of your shared inbox and use those, because real customers phrase things in ways you would never think to type.
Run the same short protocol against every tool on your shortlist. Ask a handful of your most common questions in the customer's own words. Ask one of them again in a deliberately vague paraphrase. Ask something your documentation genuinely does not cover and see whether you get a refusal and a handoff or a confident guess. Then edit one source article, change a specific detail, and ask again to see how long the change takes to show up. Finally, ask a question that requires a human, and check that the escalation lands somewhere a person will actually see it, with the conversation history attached.
Keep the transcripts. The value of this exercise is not the score, it is that you learn which of your articles are ambiguous. Most teams discover their documentation problem during evaluation and fix it regardless of which vendor they choose.
How does grounding behave on voice, WhatsApp and social channels?
Retrieval works the same way across channels, but the answer format cannot. A paragraph with two links reads fine in web chat and is useless on a phone call, where the reply has to be short, spoken in one breath, and followed by an offer to send the detail by text or email. Voice also raises a disclosure question: callers should be told they are speaking to an AI at the start, which is how the Corebee receptionist handles it, and it is worth confirming any voice vendor does the same before you point your main line at it.
Messaging channels bring a different constraint. Conversations on WhatsApp, SMS, Instagram or Telegram are long-lived and often resume days later, so grounding has to combine your documentation with the thread's own history. Ask whether the knowledge base is genuinely shared across channels or configured per channel. Separate configurations mean you will maintain the same answers more than once and they will drift apart.
Does pricing change how much grounding you actually get to use?
It does, more than most buyers expect. When AI answers are billed per resolution, every conversation the assistant handles has a marginal cost, so the rational move is to restrict it to your highest-volume questions and leave the long tail to humans. When teammates are billed per seat, you hesitate before adding the warehouse manager or the weekend contractor who occasionally needs to read a thread. Both pricing shapes quietly push you toward using less of the thing you bought.
Flat pricing removes that calculation. Corebee charges $99 per month with unlimited seats and unlimited AI conversations, no per-seat fee and no per-resolution fee, which means grounding every channel and inviting everyone who needs access costs the same as a narrow deployment. Several vendors offer flat or bundled arrangements at various tiers, so compare the total on your expected volume rather than the headline rate, and read a structured comparison such as these Zendesk alternatives before assuming the incumbent is the cheapest path.
What should a small team do first?
Start with the documentation, not the vendor. Take the questions your team answers most often, write one clear canonical answer for each, and delete the stale pages that contradict them. That work transfers to any tool you pick and it is the single biggest determinant of answer quality.
Then shortlist two tools, run the test protocol above on the same set of real conversations, and pay attention to handoff quality as closely as answer quality. Launch on one channel, read every transcript for the first stretch, and treat each bad answer as a documentation ticket rather than a verdict on the AI. Expand channels only once the answers on the first one are ones you would be comfortable sending yourself.