
How to Measure If Your Chatbot Is Actually Working
- chatbot-kpis
- chatbot-metrics
- ai-chatbots
- measurement
- customer-support-automation
- roi
Your chatbot is live. It's been answering questions for a few weeks, maybe a few months, and the dashboard your platform gives you shows conversations climbing every week. That number feels good to report in a team update. It also tells you almost nothing about whether the bot is actually helping anyone.
This is the question we hear most from founders who already built a chatbot, as opposed to the ones still deciding whether to: "how do I actually know if this thing is working?" If you're still weighing whether to build one at all, our breakdown of whether AI chatbots are worth it for startups covers that decision. If yours is built but not live yet, why most chatbot projects fail covers the pre-launch checklist. This post picks up after both of those — the bot is live, people are using it, and you need a straight answer on whether it's earning its keep.
The Quick Answer: 4 Metrics That Actually Matter
Ignore most of what's on your chatbot dashboard until you can answer these four:
- Containment rate — the share of conversations the bot resolves without pulling in a human.
- Resolution rate — of the conversations it "closes," how many actually solved the user's problem, not just ended it.
- User satisfaction on resolved conversations — how happy people are specifically with the conversations the bot handled start to finish, measured separately from escalated ones.
- Cost per conversation — total monthly chatbot cost divided by conversation volume, compared to what the human-handled equivalent costs you.
Everything else on a typical dashboard — total conversations, messages sent, average session length — is activity, not evidence.
Why "Total Conversations" Is Lying to You
Most out-of-the-box chatbot dashboards lead with the numbers that are easiest to count, not the ones that tell you anything useful. That's not a conspiracy — it's just that conversation counts and message volume are logged automatically, while "did this actually solve the person's problem" requires someone to define what a good outcome looks like and then track it. Most teams never get around to the second part.
Total conversations and messages sent
A bot that handles 3,000 conversations a month sounds like a success story until you learn that 1,800 of them ended with the user typing some version of "talk to a person" three times before giving up. Conversation volume tells you the bot got opened. It says nothing about what happened after that. Rising volume is only good news if the outcomes underneath it are also good — and you can't tell that from a single headline number.
Average session length
This one is genuinely ambiguous, which is exactly why it's a weak metric on its own. A long session might mean the bot is having a rich, useful exchange with a user working through a real question. It might also mean the user is stuck, rephrasing the same request four different ways because the bot isn't understanding them. The number looks identical in a dashboard either way. Without pairing it against resolution, you can't tell which one you're looking at.
The Four Metrics That Actually Tell You Something
Containment rate — resolved without a human
Containment rate is the percentage of conversations the bot handles entirely on its own, with no handoff to a person. It's your best signal for whether the bot is actually carrying volume rather than just fielding the first message before punting to a human.
Track it by conversation type if your platform allows it — order status, pricing questions, and appointment booking will each contain at a different rate, and blending them into one number hides where the bot is actually strong or weak. A containment rate that's dropping month over month, even while conversation volume grows, usually means the bot is encountering more variety than it was trained to handle, and it's a signal to revisit its knowledge base before the trend gets worse.
Resolution rate — solved, not just closed
This is the metric most teams skip, and it's the one that keeps containment rate honest. A conversation can "close" — the session ends, no human gets pulled in — without the person's actual problem getting solved. Someone asks about a return policy, gets a vague answer, and leaves the chat without following up. That conversation is contained. It is not resolved.
The gap between containment rate and resolution rate is the single most useful diagnostic number in this whole framework. If containment is high and resolution is meaningfully lower, your bot isn't solving problems — it's closing chat windows. That's worth catching early, because a bot that quietly closes conversations without helping people erodes trust faster than one that honestly says "I can't help with that, let me connect you with someone."
User satisfaction on resolved conversations
Most chatbot platforms let you collect a quick thumbs-up/thumbs-down or a 1-5 rating at the end of a conversation. The mistake is looking at the blended average across every conversation, resolved and escalated together. That number tells you almost nothing, because a bad rating on an escalated conversation might be about wait time on the human side, not the bot at all.
Instead, filter satisfaction to conversations the bot resolved on its own. That's the number that tells you whether the bot's actual answers are landing well with people, independent of everything else in your support process. If resolution rate looks fine but satisfaction on resolved conversations is weak, the bot is technically closing the loop but doing it in a way that leaves people annoyed — often a tone problem, an answer that's technically correct but unhelpfully vague, or a flow that makes people repeat information they already gave.
Cost per conversation
This is the number that turns "the chatbot feels like it's helping" into an actual business case. Add up what the bot costs you in a month — platform fees, any usage-based charges, and a fair share of the time your team spends maintaining it — and divide by the number of conversations it handled. Compare that to what the same volume would cost if a person handled it: your fully loaded hourly cost times the average time a human takes to answer the same kind of question.
The trap here is calculating this number using containment rate alone. If the bot "contains" 500 conversations a month but only genuinely resolves 300 of them, your real savings are based on 300, not 500 — the other 200 people either gave up unsatisfied or came back through another channel, and that cost didn't actually disappear. Cost per conversation only means something when you calculate it against resolved volume, not raw conversation count.
A Simple Monthly Review Framework
You don't need a dedicated analytics hire to keep an eye on this. A 30-minute monthly check-in covers it:
- Pull the four numbers. Containment rate, resolution rate, satisfaction on resolved conversations, and cost per conversation. Write them down somewhere you'll look at again next month — a spreadsheet is fine.
- Compare trend, not snapshot. A single month's numbers tell you less than the direction they're moving. Is containment climbing while resolution holds steady? Good. Is containment climbing while resolution slides? That's the warning sign from the section above.
- Read a handful of failed conversations. Numbers tell you something's wrong; transcripts tell you what. Pull ten conversations that ended in escalation or low satisfaction and actually read them. Patterns show up fast — the same question phrased five different ways, the same policy the bot keeps getting slightly wrong.
- Pick one thing to fix. Not five. One. Update the knowledge base entry that keeps causing bad answers, adjust the escalation trigger that's too aggressive or too passive, or rewrite the one response that reads as robotic. Ship it, then check whether next month's numbers move.
What the Numbers Are Telling You
A few common patterns, and what they usually mean:
- Low containment, high resolution — the bot is being too cautious, escalating conversations it could actually handle. Often a confidence threshold set too conservatively.
- High containment, low resolution — the bot is closing conversations without solving them. This is the one that looks good on a summary dashboard and is actually the most damaging to fix later, because it erodes trust quietly.
- Solid resolution, weak satisfaction — the bot is technically answering correctly but in a way people don't like. Look at tone, response length, and whether it's making people repeat themselves.
- Rising cost per conversation — either volume is growing faster than the bot can absorb it (more escalations creeping back in), or maintenance overhead is growing without a matching increase in what the bot handles well.
The Bottom Line
A chatbot dashboard full of green numbers doesn't mean much if the numbers aren't measuring the right things. Containment rate, resolution rate, satisfaction on resolved conversations, and cost per conversation give you an honest read on whether the bot is actually doing its job — everything else is noise dressed up as a metric.
If you're not sure your current setup is even measuring the right four numbers, that's usually a quick fix, not a rebuild. Talk to us — we'll help you get a real read on what your chatbot is doing, and a plan for the one thing worth fixing first.



