PepoChat
AnalyticsSupport operationsDefinitions

AI Support Chatbot Metrics: 9 That Matter and 3 That Mislead

Definitions, formulas and a weekly review routine for the chatbot metrics that track outcomes, plus the three vanity numbers most dashboards lead with.

PepoChat TeamPublished Last verified 13 min read
A close-up of a web analytics dashboard showing user retention by weekly cohort as a blue heat-map table

Short answer

The chatbot metrics that matter are the ones that tie an answer to an outcome: resolution rate, escalation rate, escalation quality, answer accuracy from a weekly sample, the "I don't know" rate, customer satisfaction votes, time to first human reply after a handoff, cost per conversation, and topic trends. Total conversation count, "deflection" that counts any chat that went quiet, and handling time for a bot are the three that mislead. A resolution rate is only meaningful once you have defined what "resolved" means for your own inbox.

Most chatbot dashboards are built to make the chatbot look good. They lead with the number of conversations, call every chat that ended a "deflection", and bury the question that matters, which is whether the visitor got what they came for. If you are a founder, a support lead or the one person who owns the website chat for a small team, you need a shorter list and a clearer definition of each item on it.

This post defines nine metrics for an AI support agent, gives the formula for each, says where a public benchmark exists and where it does not, and explains what to do when a number moves. It then names three metrics that look useful and are not. The examples use PepoChat, but the definitions apply to any AI chatbot that answers from a knowledge base and hands off to people.

One warning before the list. Chatbot vendors define "resolution" in ways that suit their billing, and the definitions differ from each other. The first job of a metrics review is to write down your own definition, once, and stick to it.

What is a good chatbot metric?

A good metric for an AI support agent has three properties. It measures something the visitor would recognise as an outcome, not an activity. It can be computed the same way every week from data you already have. And it points at an action: when it moves, you know which part of the system to look at, whether that is the knowledge base, the handoff rules or the team's response time.

The nine metrics below pass that test. They fall into three groups: what the agent did with the conversation, how good the answers were, and what the whole thing costs.

The nine chatbot metrics that matter

1. Resolution rate

Resolution rate is the share of conversations the AI agent closed without a person stepping in. It is sometimes called self-serve rate or containment rate, though containment is the looser term because it includes chats the visitor simply abandoned.

Formula: conversations resolved by the agent ÷ all conversations that reached the agent × 100.

The catch is the word "resolved". Zendesk's automated resolution tiers count a resolution only when the AI agent "handled the interaction to completion without the customer requesting further assistance" and a language model verifies that the request was satisfactorily resolved; conversations that fail that check are downgraded to a "contained" tier. Intercom's Fin resolutions article distinguishes a confirmed resolution, where the customer replies with something like "that helped", from an assumed one: "If a customer disengages from the conversation for 24 hours after Fin's last answer, it is considered an assumed resolution."

Those are billing definitions. For your own dashboard, pick the strictest definition you can measure. In a tool with explicit statuses, a resolved conversation is one that ended in the resolved state with no escalation; a chat the visitor walked away from is not a resolution, it is unknown. There is no universal benchmark, because the rate depends on what your visitors ask and how complete your content is; the ticket deflection benchmarks post collects the published ranges and explains why they vary so widely. Set your own baseline in the first full month and judge against that.

Worked example: 400 conversations in a month, 300 resolved by the agent, 100 escalated. Resolution rate is 300 ÷ 400 = 75%.

2. Escalation rate

Escalation rate is the share of conversations the agent handed to a person, whether because the visitor asked, the knowledge base had nothing relevant, or a rule fired.

Formula: escalated conversations ÷ all conversations × 100. In the example above, 100 ÷ 400 = 25%.

Escalation rate and resolution rate are not always mirror images, because a third bucket exists: conversations that were neither resolved nor escalated, usually because the visitor left. If the two rates do not add up to roughly 100%, the gap is your abandonment rate, and it deserves its own line.

A high escalation rate is not automatically bad. Early on it is the most useful number you have, because every escalation is a question your content did not answer. Watch the trend, not the level: it should fall for three or four weeks after launch as you fill gaps, then settle. A sudden rise usually means a new product, a new policy or a pricing change that the knowledge base has not caught up with.

3. Escalation quality

Escalation quality is whether the person who picked up an escalated conversation had what they needed to continue it without asking the visitor to start over.

This one is a judgement, not a count, so measure it by sampling. Once a week, open ten escalated conversations and answer yes or no to three questions: was the full transcript there, did the agent explain why it stopped, and did the human's first reply repeat a question the visitor had already answered? Score each conversation out of three, and track the average.

The human handoff post lists what a handoff must carry across. If the score is low, the fix is usually in the handoff design or the inbox, not in the AI.

4. Answer accuracy

Answer accuracy is the share of the agent's answers that were correct according to your own content, judged by a person who knows the product.

Formula: correct answers in the sample ÷ answers reviewed × 100.

No chatbot reports this about itself honestly, because the model cannot grade its own work. The practical method is a fixed weekly sample: pull twenty resolved conversations at random, read the agent's answers against the source pages, and mark each answer correct, partly correct or wrong. Twenty a week is small enough to sustain and large enough to catch a drift.

Treat "partly correct" as wrong for the headline number. A confident answer with one wrong detail, such as the right return window but the wrong address, costs more than an honest "I don't know", because the visitor acts on it. The hallucination guide explains why grounding the agent in retrieved content keeps this number high, and what to do when it slips.

5. "I don't know" rate

"I don't know" rate is the share of conversations in which the agent said it could not find the answer and offered a person instead of guessing.

Formula: conversations with an explicit abstention ÷ all conversations × 100.

This is the metric most teams read backwards. A rising abstention rate can be good news: it means the agent is refusing to invent answers. The number you want is not zero. What you want is for each abstention to be followed by a content fix, so that the same question does not produce another abstention next week. Pair it with the topics list: if "shipping to Canada" is the top topic and also the top abstention, that is one page to write.

When the abstention rate is near zero and answer accuracy is below 90%, the agent is guessing. That combination is the signature of a chatbot that has been tuned to sound helpful rather than to be right.

6. Customer satisfaction

Customer satisfaction, or CSAT, is the share of visitors who rated their conversation positively. Wikipedia's customer satisfaction entry defines it as the number or percentage of customers "whose reported experience with a firm, its products, or its services (ratings) exceeds specified satisfaction goals", and notes that single-item scales now perform about as well as longer surveys in commercial use.

For a chat widget, the single item is a thumbs up or thumbs down at the end of the conversation. Formula: positive votes ÷ all votes × 100.

Two cautions. First, response rates are low and skewed: people who are annoyed vote more than people who are fine, so a 70% positive share can reflect a healthy chat. Compare month to month rather than against a target you read somewhere. Second, only ask on conversations that actually ended. A prompt that appears mid-chat measures the prompt, not the chat.

7. Time to first human reply after handoff

Time to first human reply is how long a visitor waits, after the agent hands off, before a person's first message appears.

Formula: timestamp of the first operator message minus timestamp of the escalation, averaged or, better, taken at the median and the 90th percentile.

The AI answers in seconds, so the whole waiting experience now lives in this gap. A chatbot with a 75% resolution rate and a four-hour handoff wait is a worse product than one at 60% with a ten-minute wait. If you publish business hours in the widget, measure this inside hours and outside hours separately, because outside hours the honest target is "by the time we said".

8. Cost per conversation

Cost per conversation is what you pay for the chat channel, divided by the conversations it handled.

Formula: (subscription + usage charges + the share of staff time spent on escalations) ÷ conversations in the month.

Include the human time, because the point of the agent is to change that number. A flat plan makes the software part trivial to compute: on PepoChat's Starter plan at $49 a month with 2,000 replies included, a workspace that used 1,500 replies across 600 conversations pays about eight cents per conversation for the software; the pricing page has the allowances and the overage rate. On per-resolution pricing the same sum needs the vendor's definition of resolution, which is why the Intercom Fin pricing post works the example in detail.

Topic trends are the ranked list of what visitors ask about, week by week, and coverage gaps are the topics on that list that end in escalation or abstention more often than average.

There is no formula, but there is a procedure: list the top ten topics, and next to each write its escalation rate. The topics with a high share of visitors and a high escalation rate are your content backlog, in priority order. This is the metric that turns the other eight into work.

A close-up of an engraved steel ruler with millimetre and centimetre markings
A metric is only comparable when the unit stays fixed. Define 'resolved' once and do not move it.

Summary table: chatbot KPIs, formulas and cadence

MetricFormulaDirection you wantHow often to look
Resolution rateResolved by agent ÷ all conversationsUp, after accuracy is confirmedWeekly
Escalation rateEscalated ÷ all conversationsDown over the first month, then stableWeekly
Escalation qualitySampled score out of 3 on ten handoffsUpWeekly
Answer accuracyCorrect answers ÷ answers sampledUp, and never below 90%Weekly, twenty conversations
"I don't know" rateAbstentions ÷ all conversationsFalling per topic, not to zero overallWeekly
Customer satisfactionPositive votes ÷ all votesUp month to monthMonthly
Time to first human replyOperator's first message time minus escalation timeDown, at the median and 90th percentileWeekly
Cost per conversation(Software + usage + escalation staff time) ÷ conversationsDownMonthly
Topic trends and gapsTop topics ranked with each one's escalation rateGaps shrinkingWeekly

Three chatbot metrics that mislead

Total conversations

Volume is not value. A rising conversation count can mean the widget is more visible, or that a checkout bug is sending everyone to chat. Nielsen Norman Group's study of the user experience of chatbots found that "as soon as users deviated from the prescribed script, problems occurred", and concluded that improving the site itself often returns more than a chatbot that gets little use. Count conversations to size the workload; never report the number as a result.

Deflection that counts any chat that ended

Deflection is the share of contacts that never reached a person. Defined that way it silently includes visitors who gave up. Intercom's assumed resolution, where 24 hours of silence after an answer counts as a success, is the clearest example of why the number needs a stricter definition; the Fin pricing post walks through what that does to a bill, and the deflection benchmarks post shows how far published deflection figures drift from verified resolutions. If you must report deflection, report it beside answer accuracy so a reader can tell abandonment from success.

Average handling time and session length

Handling time is a human-agent metric. For a bot it is always a few seconds, so it says nothing. Session length and words per reply are worse: a long session can mean an engaged visitor or one who asked the same question four ways. If you want a speed number, use time to first human reply after handoff, which is the one wait the visitor actually feels.

A weekly 30-minute chatbot review routine

The metrics only help if someone looks at them on a schedule. This routine fits in half an hour and needs nothing more than the inbox and a spreadsheet.

  1. Five minutes: the counts. Note the week's conversations, resolved, escalated and unanswered. Compute resolution and escalation rates and add them to a running sheet.
  2. Ten minutes: the accuracy sample. Open twenty resolved conversations. Mark each answer correct or wrong against your own pages. Write down every wrong answer's topic.
  3. Five minutes: the escalations. Open ten escalated conversations. Score the handoff out of three. Note any question the agent could have answered if a page existed.
  4. Five minutes: the topics. Compare this week's top topics with last week's. Anything new that ended in escalation goes on the content list.
  5. Five minutes: the fixes. Turn the two lists into work: a wrong answer becomes a corrected page or a Q&A pair, a gap becomes a new page or help article. In PepoChat, a Q&A pair takes effect within moments and does not count as a knowledge source, so the fix is often quicker than the review.
A bearded man in a checked shirt points at orange and blue sticky notes arranged in columns on a whiteboard
Thirty minutes a week with the inbox open beats a quarterly report. Fix the content, not the symptom.

What PepoChat's analytics show, and what they do not

Being precise about what a product measures is part of being honest about metrics, so here is what you get in PepoChat as of September 2026.

The Analytics page has two cards. The Sentiments card shows the share of positive and negative votes over the last 90 days, with the vote counts; the votes are the thumbs up or down a visitor sees after a conversation is resolved, so they are real feedback, not an AI estimate of tone, and the prompt only appears on resolved chats. The Topics card shows the most asked topics over the last six weeks as a ranked list, with a weekly line for the top three; each conversation is labelled once, automatically, from the visitor's first message, with a one-to-three-word topic such as "Order Status", and labels are reused so trends are comparable. The analytics documentation describes both cards.

Usage lives on Plans & Billing rather than on Analytics: AI replies used against the allowance, knowledge sources and knowledge base text against the caps, help articles, and members against seats.

The team inbox is where the counting metrics come from. Conversations carry a status of unresolved, escalated or resolved, with filters for each, plus Mine and Unassigned for splitting the queue. Export CSV produces a file with the conversation id, status, timestamps, visitor name and email, message count and the full transcript for a date range, which is what you need for the weekly accuracy sample and for computing time to first human reply. Internal notes let a reviewer flag a conversation for a colleague without the visitor seeing it. Chats from the test drawer are excluded from analytics, from the inbox by default and from contacts, so testing does not pollute the numbers. The team inbox documentation has the details.

What is not there matters just as much. There are no charts of conversation volume, resolution rate or reply usage on the Analytics page. To get resolution and escalation rates you count from the inbox filters or from the CSV export, exactly as in the worked example: 400 rows, 300 with status resolved and no operator message, 100 escalated, gives 75% and 25%. Answer accuracy and escalation quality are always a human review, in PepoChat and everywhere else. All of this is on the free plan; the pricing page lists what changes on Starter and Growth, and it is allowances and seats, not features.

What to do next

Start with the routine, not the dashboard. Pick a definition of "resolved", run the thirty-minute review for four weeks, and you will have a baseline that means something for your own visitors. If the accuracy sample keeps finding the same kind of gap, the guide to training the agent on your website and PDFs covers how to fill it, and the best practices for accurate answers page lists the content habits that move these numbers most. To try the review on real conversations, create a free workspace; every feature described here is on the free plan.

Frequently asked questions

What are the most important chatbot metrics?
Nine cover what matters: resolution rate, escalation rate, escalation quality, answer accuracy from a weekly sample, the rate at which the agent admits it does not know, customer satisfaction votes, time to first human reply after a handoff, cost per conversation, and topic trends with their coverage gaps. Each ties an answer to an outcome and points at a fix when it moves.
How do you calculate chatbot resolution rate?
Divide the conversations the AI agent closed without a person stepping in by all conversations that reached the agent, then multiply by 100. With 400 conversations, 300 resolved by the agent and 100 escalated, the resolution rate is 75% and the escalation rate 25%. Count only conversations that ended in a resolved state; a chat the visitor abandoned is unknown, not resolved.
What is a good chatbot resolution rate?
There is no universal benchmark, because the rate depends on what your visitors ask and how complete your content is. Vendors also define resolution differently: Intercom counts 24 hours of silence after an answer as an assumed resolution, while Zendesk verifies resolutions with a language model. Set your own baseline in the first full month and judge later weeks against it.
Is a high escalation rate bad for an AI chatbot?
Not by itself. Right after launch, escalations are the most useful data you have, because each one is a question your content did not answer. The rate should fall for three or four weeks as you fill gaps, then settle. A sudden rise usually means a new product, policy or price that the knowledge base has not caught up with yet.
Which chatbot metrics are misleading?
Total conversations, because volume is not value and a spike can mean a bug elsewhere on the site. Deflection that counts any chat that went quiet, because it hides visitors who gave up. And handling time or session length for a bot, because the bot always answers in seconds and a long session can mean confusion as easily as engagement.
What analytics does PepoChat provide for its AI support agent?
A Sentiments card with the share of thumbs-up and thumbs-down votes over 90 days, a Topics card that ranks auto-labelled question topics over six weeks with a weekly line for the top three, usage meters on the billing page, inbox filters by unresolved, escalated and resolved status, and a CSV export with full transcripts. There are no resolution-rate charts; you compute those from the inbox counts or the export.

Try this on your own site in ten minutes

PepoChat includes every feature on the free plan — 200 AI replies and 5 knowledge sources a month, no credit card.