All articlesInsights

Inside a 98% resolution rate: what actually moved the number

Priya Nandakumar10 min read
Insights98%

When we compared deployments with resolution rates above 95% against the average, the model choice barely mattered. What did matter was Knowledge Base hygiene — a finding that surprised nobody once we said it out loud, but that almost nobody was actually acting on before we went looking.

We ran this comparison across a meaningful number of live deployments, specifically to find out what separated the ones performing best from the ones performing adequately, rather than assuming the answer was obvious.

Docs written for a first-timerContent Gap Analyzer, acted on weeklyOne clear answer per question

What we compared

We controlled for model, for industry, and for overall ticket volume as best we could, and still found a wide spread in resolution rate across deployments that looked, on paper, nearly identical in every other respect. That spread is what sent us looking at the Knowledge Base itself as the actual variable, rather than anything about the agent's configuration.

The top performers had documentation written for a person who's never used the product before, updated every time a policy changed, and organized so the Agent could point to one clear answer instead of piecing one together from three vague sources. The lower performers, without exception, had documentation that assumed a level of prior context a genuinely new user, or a genuinely new agent reading it fresh, didn't have.

Knowledge base hygiene

"Hygiene" is a deliberately unglamorous word for what turned out to be a genuinely high-leverage practice: a single clear source of truth per topic, phrased plainly, kept current. The deployments with the highest resolution rates treated this the way a good engineering team treats code review — someone owned it, changes to it were reviewed, and staleness was treated as a bug rather than an inevitability.

The deployments with lower resolution rates tended to have the opposite pattern: documentation written once at launch and rarely revisited, three slightly different versions of the same policy scattered across different pages, and no single owner responsible for noticing when a policy changed in the business but not yet in the docs. None of that shows up as a model problem in a support transcript. It shows up as an agent giving a technically-sourced but subtly wrong answer, because the source itself was wrong or outdated.

The Content Gap Analyzer habit

The second biggest factor was how quickly a team acted on the Content Gap Analyzer — the tool that surfaces topics where the Agent answered with low confidence repeatedly, usually because a policy exists but was never written down anywhere the Agent could retrieve it. Teams that fixed flagged gaps within a week saw resolution rates climb noticeably faster than teams that let them sit.

What separated the fast-fixing teams from the slow ones wasn't resources — some of the fastest-fixing teams were quite small. It was whether reviewing the gap report was a standing weekly habit assigned to a specific person, versus something that got looked at occasionally when someone remembered. The habit mattered more than the headcount.

None of this required more engineering. It required treating the Knowledge Base like a product surface customers actually touch — because, functionally, it is one, even though it's rarely staffed or reviewed with the same seriousness as an actual product surface would be.

Common questions

How much documentation is actually enough to get started, versus needing to be comprehensive from day one? Comprehensive from day one isn't realistic and isn't necessary. What matters more is that whatever exists is accurate and current — a smaller, well-maintained Knowledge Base outperforms a larger, stale one in every comparison we've run.

Who should actually own Knowledge Base maintenance — support, product, or someone else? The highest-performing teams we looked at had this owned by whoever is closest to policy changes as they happen, which was support leadership more often than product, simply because support tends to hear about a policy change first, through the tickets it generates.

What the gap looked like up close

One specific comparison sharpened this finding more than any other: two deployments in the same broad industry, similar size, running the same underlying model, with a resolution-rate gap of nearly fifteen points between them. Given how similar everything else was, the gap demanded an explanation, and the explanation turned out to be almost entirely about how each team treated its own documentation as an ongoing responsibility versus a one-time setup task.

The higher-performing team had a specific, named owner for Knowledge Base updates — not a committee, not "whoever notices" — and that person's role explicitly included attending the same meetings where policy changes got discussed, so documentation updates happened in parallel with the policy change itself rather than as an afterthought discovered later through a support ticket.

The lower-performing team's documentation had, by contrast, been written comprehensively at launch by a contractor, handed off, and then updated only reactively — usually after a customer had already hit a gap and a support agent had manually answered around it, without that manual answer ever making its way back into the actual Knowledge Base for the Agent to use next time. Every one of those individually-answered gaps represented a missed opportunity to close the same gap permanently.

We walked the lower-performing team through exactly what the higher-performing team's process looked like, and the change they made wasn't complicated: assign an owner, put Knowledge Base review on the agenda of the same meeting where policy changes get discussed, and treat any manually-answered gap as something that gets written back into the documentation within the same week, not eventually. Resolution rate on that team climbed roughly ten points over the following two months, without a single change to their underlying Agent configuration or model.

That result is the clearest evidence we have that this isn't really an AI story at all — it's an operations story that happens to involve AI. The lever that moved the number by the largest margin we've observed anywhere in this comparison was a staffing and process decision, not anything about model capability, prompt design, or retrieval architecture.

We've since started asking this question directly during onboarding for any new deployment: who, by name, owns Knowledge Base freshness, and is that person in the room when policy changes happen? Teams that can answer that question clearly and specifically, on day one, have consistently reached higher resolution rates faster than teams where the honest answer is "nobody specifically" or "whoever has time."

It's a small, almost boring question to ask this early in a deployment, and that's exactly why it's easy to skip — but the size of the gap it predicts is large enough that we now treat it as one of the first things worth getting right, well before tuning any model parameter or confidence threshold.

Common questions

How do you know whether a low-performing deployment's problem is Knowledge Base hygiene versus something else entirely, like a poorly-tuned confidence threshold? We check hygiene first, specifically because it's the more common root cause in our own comparisons, but a proper diagnosis looks at both — a well-maintained Knowledge Base paired with a badly-tuned threshold can still underperform, so ruling out threshold issues is part of the same diagnostic pass.

Does documentation ownership need to be a full-time role, or can it be a part-time responsibility layered onto an existing job? Part-time is fine and is in fact the more common setup among the deployments we've studied — what matters is that it's an explicit, named responsibility with enough authority to get changes made quickly, not necessarily a dedicated full-time position.

How do you measure "documentation hygiene" directly, rather than only inferring it from resolution rate after the fact? We look at a few proxy signals: average time between a policy change and the corresponding documentation update, the Content Gap Analyzer's rate of newly surfaced gaps over time, and how often the same gap reappears after supposedly being fixed — a recurring gap usually means the fix didn't actually address the root cause.

Is there a point at which more documentation stops helping and starts actively hurting resolution rate, through sheer volume or redundancy? Yes, and this is a real failure mode distinct from staleness — a Knowledge Base with many overlapping, redundant documents on the same topic can confuse retrieval the same way outdated documents do, which is part of why we recommend periodic consolidation, not just periodic freshness checks.

What's a reasonable timeline for a team starting from a genuinely weak Knowledge Base to reach a resolution rate comparable to the top-performing deployments described here? Based on the cases we've walked through, a few months of consistent weekly attention is usually enough to close most of the gap, assuming the team commits to the ownership and review habits described above rather than treating the initial cleanup as a one-time project.

Is there a risk of over-investing in documentation at the expense of other improvements? We haven't seen this in practice — the marginal hour spent on documentation hygiene has consistently outperformed the marginal hour spent on most other levers we've measured, for teams below the top resolution-rate tier. The returns diminish eventually, but most teams are nowhere near that point yet.

How do you know if your Knowledge Base has a hygiene problem before it shows up in resolution rate? The Content Gap Analyzer is the earliest signal — a rising rate of low-confidence answers on a specific topic almost always means the underlying documentation for that topic has drifted out of date, well before it's visible in a top-line resolution number.

It's also worth naming what this finding doesn't mean: it's not a claim that model quality is irrelevant, only that it explained less of the variance between our best and average deployments than documentation hygiene did, in the specific comparisons we ran. A sufficiently weak model paired with excellent documentation would still underperform a strong model paired with excellent documentation — the finding is about where the marginal improvement was actually coming from at the resolution-rate levels these deployments were already operating at, not a universal ranking of every possible lever.

We'd encourage any team running a similar comparison internally to be equally careful about that distinction, because it's easy to overcorrect from "documentation mattered more than we expected" to "documentation is the only thing that matters," and the second claim isn't one our data actually supports.

One more distinction worth drawing out: "documentation hygiene" as we've described it here is not the same thing as documentation volume, and conflating the two is a mistake we've watched more than one team make after first hearing this finding. A team that responds to this piece by writing dramatically more documentation, without also building the ownership and update-cadence habits described above, is likely to end up with a larger Knowledge Base that's just as prone to drifting stale as a smaller one — more surface area to maintain, without the maintenance discipline that actually determines whether that surface area stays accurate.

The teams we've seen succeed at this took the opposite instinct: several actually reduced their total document count during the cleanup process described earlier, consolidating three overlapping, slightly-inconsistent documents on the same topic into one clear, owned source — smaller, not larger, and meaningfully easier to keep current as a direct result.

We'll close with the specific number that made this finding stick internally: the gap between our best and average deployments' resolution rates was wider than the gap between our best and worst model choices across the same deployments, measured independently. That comparison, more than any single case study, is what convinced our own team to stop treating model selection as the primary lever and start treating documentation ownership as at least an equal priority in every new deployment's onboarding checklist.

We've since made this specific comparison — documentation-hygiene variance versus model-choice variance — a standing slide in how we onboard new team members, precisely because it's the single fastest way we've found to correct the assumption that a better model is always the right first thing to reach for when a deployment's resolution rate isn't where anyone wants it to be.