AI in replies

A weekly quality review for your AI replies

By Iryna Savchenko 3 min read
A weekly quality review for your AI replies

The quality of AI replies isn’t measured by a platform dashboard but by reading conversations — fifteen minutes a week and two specific slices: threads that went silent, and threads a human had to take over. That’s where you see the gaps.

The main trap is looking at reply volume. An assistant can answer a hundred messages and produce no orders; it looks like work and isn’t.

A laptop showing a customer chat next to a notebook on a desk
Photo: https://kaboompics.com/ / Pexels

Which conversations to read and which to skip

Only two types are worth reading. The rest will tell you nothing.

  1. 1

    Abandoned threads

    The assistant replied, the customer went quiet. Usual cause: the reply offered no next step.

  2. 2

    Handed to a manager

    This is where the automation boundary shows: either it correctly passed something complex, or it failed something simple.

  3. 3

    Repeated questions

    The customer asked the same thing twice — the answer was technically correct and unclear.

  4. 4

    Successful ones — a sample

    Two or three, to check tone. No more: they confirm rather than teach.

What exactly to measure

Three numbers you can count by hand on your sample:

  1. How many threads reached a next step — the customer named a size, a city, or agreed to order. That’s the actual value.
  2. How many went to a human — and crucially, why. A correct handover (complaint, haggling) is a success, not a failure. A handover because it «didn’t understand the question» is a failure.
  3. How many went silent after the assistant’s reply. If that share doesn’t fall week over week, the problem is in the structure of your replies rather than individual phrases.

Fix the instruction, or add a scenario?

The rule that saves months: a scenario cures one case, an instruction cures a class of cases. If the mistake could recur in a different topic, it belongs in the instructions.

An example. The assistant invented a delivery time for a city. The weak fix is a scenario: «that city → 2 days». The strong fix is a rule: «quote delivery times only from the shipping table; for cities not in the table, say you’ll check and get back». The second closes every city at once, including the ones nobody has asked about yet.

How not to break what already works

One change per iteration. It’s boring and it’s the only way to know what did the work.

Before editing, save a handful of real questions as a test set: stock, price, delivery, haggling, complaint. After the edit, run all of them. If the price answer improved and the complaint handover broke, you’ll see it immediately instead of hearing it from a customer next week.

Weekly review checklist
  1. Open the week’s conversations; filter for abandoned and handed-over threads.
  2. Read 10–15 and note recurring causes — patterns, not individual mistakes.
  3. Count three numbers: reached a step / went to a human / went silent.
  4. Pick ONE instruction change that closes the most frequent pattern.
  5. Run the test question set before and after the change.
  6. Write one row into the weekly table.

Frequently asked questions

How many conversations should you read each week?

Ten to fifteen is enough if you pick them properly: not random ones, but the threads that went silent after the assistant replied and the ones a manager had to take over. That's where the assistant's limits show.

What's the main quality metric for AI replies?

The share of conversations the assistant carried to a concrete next step — without human help and without the customer repeating themselves. Reply volume says nothing: a bot can answer a hundred times and sell nothing.

What do you do when the assistant gets something wrong?

First work out what it was missing: data, permission to say «I don't know», or a clear prohibition. Fix the instruction, not the scenario — a new scenario cures one case, a rule cures a class of cases.

How often should you change the instructions?

No more than once a week, and one change at a time. Several edits at once make it impossible to tell which one helped and which one broke something.

How do you know AI simply doesn't fit here?

When after several iterations the same category of question still ends up with a human. That's not a failure: it means those conversations need a decision, not an answer. Take them out of automation deliberately.