Why does an AI support agent invent product details in the first place?
A language model produces the most plausible continuation of a conversation. When the correct passage is in front of it, plausible and correct are the same thing. When the passage is missing, plausible is all that remains, and the answer comes out fluent, confident and wrong. That is the entire mechanism, and understanding it changes where you look for a fix.
In practice, invented product details trace back to a short list of ordinary causes:
- The answer was never written down anywhere the agent can read. Someone on the team knows the shipping cutoff for the warehouse, but it lives in their head.
- The answer exists in a format the agent cannot parse cleanly. A spec sheet locked inside a PDF image, a returns policy pasted as a screenshot, a compatibility table with merged cells.
- The answer exists twice, in two versions, and the agent picked the stale one. An old landing page still says the older model ships with the cable; the new one does not.
- The question is about something the business genuinely has not decided, such as whether a discontinued size is returning.
- The agent was configured to be helpful above all else, with nothing in its instructions telling it that declining to answer is an acceptable, even desirable, outcome.
Only the last of those is about the software, and none of them are about the model being untrustworthy in some abstract sense. Your content and your settings are the levers, which is why the rest of this post is mostly about writing and reviewing rather than about prompts.
How do you give the agent a source it can actually quote?
A grounded agent searches your material, retrieves the passages that look relevant, and writes its answer from those passages. Everything about the quality of that answer depends on whether the right passage exists and whether the retrieval step can find it. So write for retrieval, not for a brochure.
Start from what customers actually ask. Go through your past tickets and chat transcripts and pull the recurring questions in the customer's own words, including the awkward phrasings and the misspellings of your product names. That list, not your internal product taxonomy, is the outline for your knowledge base.
Then apply a few rules that make a real difference:
One topic per article, with the question as the heading. An article titled "Do you ship to Canada?" retrieves far better than a section buried inside "Shipping and Delivery Information" that covers domestic, international, expedited and freight in one long page. Long mixed pages get chunked into fragments that each look partly relevant and fully relevant to nothing.
One canonical home for every fact. Prices, warranty terms, return windows, lead times and compatibility should live in exactly one article, with other pages linking to it rather than restating it. Duplicated facts are how stale answers survive; the agent has no way to know which copy you updated.
Specifics in text, never only in images. Dimensions, materials, voltages, size charts, part numbers. If it is only in a photograph or a design asset, the agent cannot read it and will reach for something that sounds right instead.
Write the boundaries down too. "We do not offer installation" and "this model is not compatible with the previous generation bracket" are answers. Negative facts prevent more hallucinations than positive ones, because they cover exactly the territory where an agent would otherwise improvise.
A knowledge base built this way does double duty: customers who prefer self-service find the article, and the agent has something solid to quote when they do not.
Should the agent be allowed to answer when it is not sure?
No, and this is the single configuration choice that separates a support agent you can leave running from one you have to babysit. An agent that always produces an answer will produce a wrong answer whenever the source is thin. An agent that can decline will simply hand the thin cases to a person.
So design for refusal explicitly. The instructions should tell the agent to answer only from retrieved material, to name what it does not know rather than filling the gap, and to offer a handoff whenever confidence is low or the customer asks for a human. The customer experience of "I am not certain about that one, let me bring in a colleague" is good. The customer experience of a confident wrong shipping date is a refund request and a bad review.
Corebee follows that pattern: it answers from your website and knowledge base, and it hands off to a human when the customer asks or when it is unsure, with the conversation and its history landing in a shared inbox so whoever picks it up is not starting cold. Whatever tool you use, check that handoff is a first class behavior and not an afterthought bolted onto a chat widget.
One related point on voice. If you run an AI receptionist on your phone line, it should tell callers it is an AI. Callers who know they are speaking to software ask simpler, clearer questions and are far more willing to be transferred, which cuts the pressure on the agent to bluff.
Which questions should never be answered by the model at all?
Some questions should bypass generation entirely, because no amount of good writing makes a static document current. Route these to a live source or to a human:
- Order status, tracking and delivery dates. These come from your order system through an integration, or they do not get answered.
- Stock levels and back in stock timing. Inventory changes faster than any knowledge base.
- Account specific balances, credits and entitlements. If the agent cannot read the account, it must not describe it.
- Anything with legal or safety weight. Warranty determinations, medical or financial guidance, regulatory claims. Route to a person by policy, not by confidence score.
- Custom quotes and negotiated terms. A booking is the correct answer here, not a number.
For that last category, the useful behavior is to move the conversation forward without inventing anything: book a meeting through your calendar tool and push the contact into your CRM so a human picks it up with context. Corebee books through Cal.com or Calendly and syncs to HubSpot, Salesforce, Pipedrive, monday and Zoho, which turns "I cannot answer that" into a scheduled call rather than a dead end.
How do you test for hallucinations before customers see them?
Treat it as a test suite you run, not a feeling you develop. Build a question set and keep it in a spreadsheet. It should include:
- Real questions lifted from past tickets, with the answers you know are correct.
- Questions about products or services you do not offer, to see whether the agent invents them.
- Questions with a false premise built in, such as asking when the lifetime warranty begins on a product that has a limited warranty. A grounded agent corrects the premise; a bluffing agent runs with it.
- Compound questions that bundle a policy question with a product question, since these are where agents tend to answer one half well and improvise the other.
- Questions phrased the way customers actually type them, including in other languages if you serve them.
- Questions about competitors and comparisons, which invite confident invention more than almost anything else.
Grade each response into three buckets: correct and supported by a source, correctly declined or handed off, or wrong. Only the third bucket is a failure. When you find one, resist the urge to patch the prompt with a special case, because a prompt full of exceptions becomes unmaintainable quickly. Ask instead which article was missing, ambiguous or stale, fix the content, and re run the set.
Run the set on every channel you have turned on. Formatting and message length differ between web chat, email, WhatsApp, SMS and voice, and an answer that reads well in a chat bubble can lose its caveats when squeezed into a spoken reply.
What do you review once the agent is live?
Testing catches the failures you thought of. Live review catches the ones you did not. Put a recurring slot in one person's calendar to read a sample of real conversations, and give that person the authority to edit the knowledge base directly. A review loop that requires a ticket to marketing dies within a month.
When reading, look for a specific tell: any concrete detail in an answer that you cannot immediately point to in a source article. Measurements, timeframes, compatibility claims, policy carve outs. Those are where invention hides, and they read perfectly naturally, which is why skimming does not catch them.
Also read the handoffs, not just the answers. Every handoff is a labeled gap in your content, sorted by how often customers hit it. Cluster them, write the missing articles, and watch the same cluster shrink next month. Keep a short log of what you changed and when, so that when an answer changes you know which articles need to change with it.
How do the main platforms differ on grounding?
Every serious vendor now retrieves from your content rather than answering from general knowledge, so the differences are in control, in where live data comes from, and in what the pricing encourages you to do.
Zendesk and Freshdesk are built around mature ticketing and help center products, which means the knowledge base discipline described above is well supported and familiar to larger support teams. Intercom's Fin gives you fine grained control over which content sources are in scope and how the agent behaves when it cannot find an answer. Gorgias is oriented toward ecommerce, with tight access to order data, which matters because order status is exactly the category that should never be generated. Help Scout keeps docs and inbox simple and readable, which suits small teams who will actually maintain the content. Chatwoot is open source and self hostable, a good fit if you need the data to stay on your infrastructure. Tidio pairs live chat with an AI agent aimed at smaller online stores.
The part worth thinking about carefully is pricing, because it shapes behavior. Per seat pricing discourages you from adding the colleague who would review transcripts. Per resolution pricing puts a meter on the adversarial testing described above, and quietly rewards an agent that closes conversations rather than one that hands off honestly. Flat pricing removes that tension, which is the reasoning behind Corebee's flat $99 per month with unlimited seats and unlimited AI conversations: you can put every teammate in the inbox and run your question set as often as you like without doing arithmetic first. If you are comparing against an incumbent, the Zendesk alternatives breakdown covers the tradeoffs in more detail.
What does a practical first month look like?
Week one, pull the recurring questions out of your inbox and write short, question titled articles for the ones with stable answers. Week two, connect the agent, turn on a single channel, set the instructions so that declining and handing off are allowed and expected, and route order status and account questions to a human or an integration. Week three, run your question set, grade the answers, and fix content rather than prompts. Week four, read a sample of live conversations, close the gaps the handoffs exposed, and only then turn on the next channel.
Setup on a modern tool is not the hard part. Corebee takes about ten minutes and no developer, and most competitors are in the same range for a basic install. The work that determines whether your agent invents product details is the writing and the reviewing, and that work does not have a shortcut.