Everyone has met a bad chatbot. It cannot answer the question, will not let you reach a person, and asks you to rephrase for the third time. Those bots set a low bar. Clearing it is not hard; building an agent people actually prefer to the alternative takes discipline.
Start with the job, not the channel
Decide what the agent is for before deciding where it lives. Booking an appointment, qualifying a lead, answering order-status questions and handling returns are different jobs with different success metrics. An agent scoped to 'answer questions about our business' is scoped to fail.
Pick one or two high-volume, well-defined jobs. Do them properly. Expand once the data shows it is working.
Escalation is a feature
The fastest way to destroy trust is to trap someone in a loop. Design the exit first: a clear route to a human, triggered automatically after repeated failure, available on request at any point, and carrying the full conversation so nobody has to repeat themselves.
On a healthcare build, a single rule (two failed turns and you reach a person) removed the biggest complaint about the previous system and cut abandonment sharply. The agent handled fewer conversations end to end, and satisfaction went up.
Grounding beats cleverness
An agent that answers from your real policies, catalogue and account data is useful. One improvising from general knowledge is dangerous. Connect it to the systems of record, retrieve before responding, and prefer 'I do not have that information, let me get someone who does' to a confident guess.
- Retrieve from approved sources only, and cite them where useful
- Use tool calls for anything transactional, bookings, refunds, order lookups
- Validate every action against business rules before executing it
- Log every action with the reasoning, for audit and debugging
Voice is its own discipline
Voice agents fail on things text never exposes. Latency above about a second feels broken. Interruptions have to be handled gracefully. Numbers, names and addresses need confirmation strategies. And people speak differently from how they type, more loosely, with more hesitation.
Tune for turn-taking and latency before tuning for eloquence. A voice agent that responds quickly and plainly beats one that responds beautifully after a pause.
Measure outcomes, not conversations
Conversation volume is a vanity metric. Measure what the business cares about:
- Resolution rate, handled end to end without a human, and correctly
- Conversion, bookings made, leads qualified, sales completed
- Escalation quality, did the human receive enough context to finish quickly?
- Satisfaction, on handled conversations, compared with the human baseline
- Containment honesty, conversations abandoned are failures, not deflections
Evaluate before every change
Prompt tweaks and model upgrades change behaviour in ways that are hard to predict. Keep an evaluation set of real conversations scored against an agreed rubric, and run it before anything ships. Review a sample of live conversations every week as well, that is where you find the failure modes nobody thought to test.
What good looks like in practice
A well-built agent handles the majority of routine contacts end to end, escalates the rest cleanly, and scores at or above the human baseline on the conversations it handles. It does not pretend to be a person, and it does not attempt things it is not scoped for.
If you are considering chat or voice agents and want to know what is realistic for your volume and use case, get in touch.



