Why we started measuring trust, not just resolution rate
Resolution rate answers one question: did the conversation end without a human being pulled in? It says nothing about whether the customer left annoyed, confused, or quietly planning to churn — and that gap is why we added a second metric that, on its own, sounds almost impossible to measure well.
What resolution rate misses
A conversation can end with the Agent successfully answering every question asked, technically resolving the ticket without escalation, while the customer walks away irritated by how long it took, how many times they had to repeat themselves, or how impersonal the exchange felt. None of that shows up in resolution rate, which only tracks whether a human got pulled in — a customer who resolves their issue but leaves frustrated looks identical, on that metric, to one who resolves it and leaves satisfied.
We suspected this gap existed before we had a way to measure it, mostly from anecdotal complaints that didn't line up with what our resolution-rate numbers were telling us. Building a second metric specifically to check that suspicion was the only way to find out whether it was a real, systematic pattern or just a handful of loud outliers.
Measuring tone instead
We started sampling closed conversations and running them through a second model whose only job is to guess, from the transcript alone, whether the customer's tone shifted positive, neutral, or negative by the last message. The model doing this tone assessment is deliberately separate from the Agent handling the conversation itself, specifically so it isn't influenced by whatever the Agent itself concluded about how the conversation went — an independent read matters more than a self-reported one here.
This second model isn't trying to determine whether the ticket was resolved correctly in a factual sense — that's what resolution rate and accuracy metrics already cover. It's specifically trying to catch the more human, harder-to-quantify signal of whether the interaction itself felt good to be on the other end of, regardless of whether the underlying facts were handled correctly.
Where the two disagree
The two metrics disagree more than we expected — some "resolved" tickets read as clearly frustrated once you look at tone instead of outcome. Those are the ones we now review first, specifically because the disagreement itself is the most informative signal: a ticket where both metrics agree, whether both positive or both negative, is less interesting to investigate than one where the ticket counted as a clean resolution on paper but the tone read reflects real, unresolved frustration.
The pattern behind most of those disagreements, once we started reviewing them, tended to be conversations that took longer than they should have to reach a correct answer — technically resolved, but only after enough back-and-forth that the customer's patience had visibly worn down by the final message. Resolution rate alone would never have surfaced that as a problem worth fixing, since the end state still counted as a success.
Resolution rate still matters for capacity planning — it's still the right number for understanding how much human support capacity a given ticket volume actually requires. But trust, measured this way, is what actually predicts whether a customer sticks around after their first AI conversation, which is a different, and in some ways more important, question for the long-term health of the relationship with that customer.
Common questions
How do you validate that the tone-reading model is actually accurate, rather than just confidently guessing at sentiment? We spot-check its assessments against human review on a sample, the same way we'd validate any other model-based metric, and have tuned it specifically against cases where we have independent confirmation of customer sentiment — a follow-up survey response, a support escalation initiated later, that kind of external signal.
Does this metric get gamed if the Agent's language is tuned to sound cheerful regardless of actual outcome? This is a real risk we watch for specifically — a model that's learned to produce cheerful-sounding language independent of actual resolution quality would show a misleadingly good trust score, which is part of why we cross-reference tone against resolution outcome rather than trusting either number in isolation.
A batch of disagreement cases, reviewed closely
To understand the disagreement between resolution rate and tone-based trust scoring concretely, we pulled a batch of thirty conversations where the two metrics disagreed most sharply — technically resolved, but reading as clearly negative on tone — and reviewed every one manually, specifically looking for a common thread.
More than half of that batch shared a specific pattern: the correct answer had been given, but only after the customer had to restate their question, or a key piece of context, more than once. In several cases, this tracked directly back to the same conversation-state issue we've described elsewhere in this series — a follow-up question interrupted the Agent's tracking of the original context, and the customer's evident frustration in their later messages was a direct, legible response to having to re-explain something they'd already said.
A smaller but notable share of the batch involved conversations that were technically resolved correctly on the first attempt, but where the tone shift toward negative had started before the Agent said anything at all — customers arriving already frustrated, from a bad experience with the product itself rather than with the support interaction. Those cases are a useful, if slightly uncomfortable, finding: the tone-based metric doesn't only measure the quality of the Agent's handling, it also partly measures the customer's starting emotional state, which the Agent had no ability to change no matter how well it performed.
We had to build a rough adjustment for this second pattern specifically — cases where negative tone was clearly present from the customer's very first message — to avoid unfairly attributing pre-existing frustration to the Agent's performance. The adjusted metric now weights tone shift over the course of the conversation more heavily than the absolute tone of the final message, which better isolates what the Agent actually contributed versus what the customer brought in with them.
The remaining cases in the batch, a smaller fraction, genuinely didn't have an obvious explanation even after close review — a reminder that not every disagreement between the two metrics points to a fixable process issue, and that some variance in how a specific conversation lands with a specific customer is simply outside what any metric, however well-designed, can fully explain.
What this closer review changed, more than any specific fix, was our own confidence in treating the two metrics as complementary rather than redundant. Resolution rate alone would have called all thirty of these conversations a success. The tone-based metric alone, without the review that separated pre-existing frustration from Agent-caused frustration, would have overstated how much of the negative signal was actually the Agent's doing. Looking at both, and at the actual transcripts behind the disagreement, is what let us find the real, fixable pattern hiding inside a larger, noisier signal.
Common questions
Does the tone-based trust metric get run on every conversation, or only a sample, and how does that affect what conclusions you can draw from it? It runs on every closed conversation automatically, since the assessment itself is cheap to compute at scale — sampling comes in only for the deeper human review of disagreement cases, which is the more expensive step that doesn't need to run on every single conversation to be useful.
Could a customer's tone shift negative for reasons that have nothing to do with the specific conversation at all, like an unrelated bad day, and does that get filtered out anywhere? This is a real source of noise the metric can't fully filter out, and it's part of why we look at aggregate patterns across many conversations on a similar topic rather than treating any single conversation's tone score as a definitive individual judgment.
How do you validate that the separate tone-assessment model itself isn't just reflecting its own biases about what "sounds negative," rather than accurately reading actual customer sentiment? We validate against independent signals where we have them — a customer who later files a complaint or a churn event correlates with prior tone scores, which gives us external confirmation beyond just trusting the model's own self-consistency.
Does a consistently high trust score ever mask a resolution-rate problem, the same way a high resolution rate can mask a trust problem? It's possible in principle, though we haven't seen this pattern show up as strongly in practice — a conversation that goes badly on substance (wrong or incomplete resolution) tends to also show up as tone-negative, since customers generally respond to a bad outcome with visible frustration, making this particular blind spot less common than the reverse one this piece focuses on.
How actionable is this metric on its own, versus needing the kind of manual review described in this piece to actually fix anything? The metric alone tells you where to look; it doesn't tell you what to fix. Every meaningful improvement we've made because of this metric has required the kind of close manual review of specific disagreement cases described above — the aggregate number is a pointer, not a diagnosis.
How often do you actually sample conversations for this, versus reviewing every single one? We sample rather than reviewing every conversation, since the tone-assessment model itself runs on every closed ticket automatically, but human review of the disagreement cases happens on a sampled, prioritized basis rather than exhaustively, simply because the volume of full agreement cases doesn't need the same level of scrutiny.
Has this metric changed anything about how the Agent itself is designed, beyond just being a measurement? Yes — the state-tracking discipline we've described elsewhere in this series, specifically around not losing context across an interruption, came directly out of investigating a batch of disagreement cases and finding that a disproportionate share of them involved exactly that kind of context loss mid-conversation.
The single change we'd recommend to any team only tracking resolution rate today is not to replace it, but to sit a second, independent signal next to it and specifically look for the cases where the two disagree — because, as this piece has tried to show, the disagreement cases are where the actual, fixable problems tend to hide, far more often than in the cases where both metrics already agree with each other.
It's worth being explicit that building this second metric took real, dedicated engineering effort — a separate model, a separate evaluation process, its own periodic validation against external signals — and wasn't free in the way that simply looking at resolution rate more carefully would have been. We'd still make the same investment again, but we think it's worth naming that cost plainly rather than implying this kind of measurement is something any team can bolt on trivially; the value described throughout this piece came from taking the measurement seriously enough to validate it properly, not from adding a rough sentiment score as an afterthought.
The single sentence we'd want a support leader to remember from this piece, if nothing else sticks: a ticket counted as resolved and a customer who left satisfied are correlated, not identical, and the gap between those two things is exactly where a second metric, built carefully enough to trust, earns its cost.
We'd encourage any team building something similar to resist the temptation to reduce this to a single blended "quality" score that folds resolution and tone together — keeping them visibly separate, even though it's less tidy on a dashboard, is exactly what let us find the specific disagreement cases that turned out to matter most, and folding them together would have hidden precisely the signal this whole piece has been about surfacing.