Chatbots have a poor reputation, earned honestly over a decade of deployments that made things worse. The bad ones are still common enough that many buyers assume the category is the problem.
It is worth being specific about what a good one does, because the two jobs it can do are quite different and confusing them is the usual cause of failure.
Two jobs, two measures
An acquisition bot exists to capture enquiries that would otherwise leave — the visitor with a question at 9pm who is not going to fill in a contact form and wait until Monday. Success is qualified enquiries captured.
A support bot exists to resolve documented questions before they become tickets. Success is questions genuinely resolved.
These need different placements, different tone and different escalation rules. Deploying one bot to do both, measured on a blended number, is why so many deployments satisfy nobody.
Where the value actually concentrates
Not during working hours. If your team answers quickly, a good human beats a good bot on nearly every measure during the day.
The value is in evenings, weekends and timezone gaps — the Friday 9pm enquiry that would otherwise sit until Monday, by which point two competitors have already replied.
Scoping the deployment to that window keeps the failure surface small and makes the economics obvious. A bot that handles the overnight gap is easy to justify; a bot that intercepts every daytime visitor is harder to defend and more likely to irritate.
Check your enquiry timestamps before deploying. If nothing arrives outside working hours, the acquisition case is weaker than it looks and support deflection may be the better job.
Grounding beats model choice
The most common technical question is which model to use. It matters far less than what the model is allowed to draw on.
A bot grounded in your actual documentation — service pages, pricing, FAQs, help articles — answers accurately within a known boundary. An ungrounded bot improvises, and improvisation about your pricing or your contract terms is a liability rather than a feature.
This also means the quality ceiling is set by your documentation. Teams frequently discover during deployment that their own answers are incomplete, which is useful information regardless of the bot.
How bad deployments increase support load
The failure pattern is consistent. The bot cannot answer, does not offer a route to a person, and the visitor either leaves or arrives in the support queue frustrated and having already explained the problem once.
That produces a metric that looks like success — high containment, fewer immediate tickets — attached to a worse experience and, frequently, more work per ticket.
The difference between the two outcomes
The left column is what most people have experienced as a customer.
Adds friction
- Human route hidden or absent
- Loops when it cannot answer
- Improvises on pricing and terms
- Pretends to be a named person
- Measured on deflection alone
- Deployed sitewide on day one
Removes friction
- Human route visible from message one
- Hands off with full transcript attached
- Grounded in documented answers only
- States plainly that it is an agent
- Measured on genuine resolution
- Deployed on high-intent pages first
A practical rollout
Start on two or three high-intent pages — pricing, a primary service page, contact. Not sitewide. This limits the blast radius and concentrates the bot where questions are commercially relevant.
Ground it in those pages plus your FAQs. Set explicit refusal boundaries around pricing beyond published figures, contract terms and anything legal.
Then read transcripts weekly for the first month. Every deployment we have seen needed scope adjustments in that period, and every one of those adjustments came from reading a conversation rather than from a dashboard.
What to measure
- Qualified enquiries captured outside working hours — the acquisition case.
- Genuine resolution rate, verified by whether the person returned via another channel.
- Handoff rate and handoff quality, not containment alone.
- Conversion of bot-touched sessions versus comparable sessions.
- Escalation satisfaction — how people feel about the conversation they were handed.