All articlesInsights

The Content Gap Analyzer: finding what your Knowledge Base is missing

Priya Nandakumar10 min read
Insights

Most teams don't know what their AI Agent can't answer until a customer hits the gap. Here's how we surface it earlier, and why the tool that does this ended up being one of the most consistently useful pieces of the whole platform, despite being one of the least glamorous.

How gaps get found

?Refund window for gift cards??Do you ship to APO addresses??Can I pause a subscription?

Most teams don't find out their Agent can't answer something until a customer hits the gap and gets a bad answer or an unnecessary handoff. The Content Gap Analyzer flags those gaps before that happens, by clustering unanswered or low-confidence questions — grouping together conversations where the Agent's confidence score fell below threshold, and surfacing the underlying topic those low-confidence moments share, rather than presenting each low-confidence conversation as an isolated incident.

The clustering step is what makes this genuinely useful rather than just a raw log of uncertain moments. A handful of individually low-confidence conversations about slightly different specific questions can look like noise in a raw log; clustered together, they often reveal a single underlying gap — a policy that exists but was never written down anywhere retrievable, or a recent product change the documentation hasn't caught up with yet.

What a gap actually looks like

Every week it surfaces the topics where the Agent answered with low confidence three or more times — often revealing that a policy exists but was never written down anywhere the Agent could retrieve it. A concrete example from one deployment: a cluster of low-confidence answers about a specific warranty exception turned out to trace back to a verbal policy the support team had been applying consistently for months, that had simply never made it into written documentation anywhere the Agent's Knowledge Base could see.

That kind of gap is invisible to almost any other kind of review, because the human team already knows the policy — it's only invisible to the Agent, and only shows up as a pattern once enough low-confidence conversations about the same underlying topic accumulate to be clustered together. A single instance would have looked like an edge case. Three or four in a week is a pattern worth fixing.

Turning it into a habit

Teams that review this weekly close the top three gaps in under an hour, because the fix is usually a single missing paragraph, not a new integration or model change. The fix being small is actually part of why this needs to be a habit rather than a one-time audit — small fixes are easy to defer indefinitely if there's no standing weekly cadence forcing the review to happen, precisely because none of them individually feel urgent.

The teams with the highest resolution rates aren't the ones with the most documentation — they're the ones who treat this report as a to-do list instead of a dashboard. That distinction between "dashboard to glance at" and "to-do list to actually clear" turned out to predict resolution-rate improvement better than almost any other single behavior we could measure across deployments.

We've watched teams go from letting gaps accumulate for a month to reviewing weekly, and the resolution-rate improvement shows up within a few weeks of making that switch, without any change to the underlying model or Agent configuration — the entire improvement comes from closing gaps faster after they're already being detected, which says something about how much low-hanging fruit tends to sit undetected in most Knowledge Bases even after the initial launch effort.

Common questions

How is a "gap" different from a question the Agent just got wrong due to a model limitation? A gap specifically means the Agent lacked a retrievable source to answer from — it's a documentation problem, not a reasoning problem, and the Content Gap Analyzer is built to isolate exactly that category rather than surfacing every kind of imperfect answer.

Does the tool ever flag something as a gap when the real problem is that the question itself was unusually rare or unclear? Occasionally, and that's a normal part of using the report — not every cluster represents a genuine, worth-fixing gap, and part of the weekly review habit is exercising judgment about which clusters represent a real documentation problem versus an unusual one-off.

A specific gap, start to finish

One deployment's experience with a specific gap is worth walking through in detail, because it shows the full lifecycle of how the tool is meant to work, from first detection through to a fix and its measured effect.

The cluster first appeared as four low-confidence conversations in a single week, all loosely about whether a specific loyalty-program discount could be combined with a separate promotional code. None of the four questions were worded identically, which is exactly why they hadn't been noticed as a pattern by anyone manually skimming transcripts — each one, read individually, looked like a slightly unusual one-off question rather than part of a recognizable cluster.

The weekly review, looking specifically at the clustered report rather than individual transcripts, flagged this immediately as a candidate gap. Checking with the team that owned the loyalty program confirmed there was, in fact, a specific rule about stacking — loyalty discounts and promotional codes couldn't be combined, full stop — but that rule had never been written into any documentation the Agent's Knowledge Base could retrieve; it existed only as an assumption inside the team that built the loyalty program originally.

Writing the fix took about fifteen minutes: a single new paragraph added to the existing discounts documentation, stating the stacking rule plainly and specifically enough to be unambiguous. The following week's Content Gap Analyzer report showed zero new low-confidence instances on this specific topic, and a spot-check of a few subsequent real conversations about the same question showed the Agent answering directly and correctly, citing the newly added paragraph as its source.

What makes this example useful beyond its own specifics is how ordinary it is. Nothing about this particular gap was unusual or hard to fix once found — the entire lifecycle, from four confusing-looking individual questions to a clear pattern to a fifteen-minute fix, took about a week, almost entirely because someone was looking at the clustered report on a fixed weekly schedule rather than waiting for the pattern to become obvious some other way.

We've collected enough of these individually-small, collectively-significant examples across different deployments to be confident the pattern generalizes: the vast majority of Knowledge Base gaps are not hard to fix once identified. The entire difficulty is in identification, at a stage early enough that the gap hasn't yet generated a meaningful number of bad customer experiences — which is precisely the job the clustering and the weekly habit are built to do together, neither one being sufficient without the other.

Common questions

How does the tool distinguish a genuine documentation gap from a question that's simply outside the scope of what the Agent should ever be expected to answer? This distinction is ultimately a human judgment call made during the weekly review, not something the clustering algorithm decides on its own — the tool's job is surfacing candidate clusters, and part of the review habit is exercising judgment about which clusters represent a real, worth-fixing gap.

Does the tool work as well for a brand-new deployment with very little conversation history, or does it need volume to be useful? It needs some minimum volume to form meaningful clusters — a handful of conversations isn't enough to distinguish a real pattern from coincidence — so newer deployments typically don't get much value from it until they've accumulated a few weeks of real traffic.

Can a gap get flagged, fixed, and then reappear later if the underlying policy changes again without anyone updating the fix? Yes, and this is exactly the kind of drift the weekly review habit is meant to catch quickly rather than letting it recur silently — a gap that was fixed once and reappears is treated with the same urgency as a brand-new gap, not deprioritized because "it was already handled before."

How do you prioritize which flagged gaps to fix first when several show up in the same week? We prioritize by estimated impact — how many real conversations the gap is likely affecting, based on cluster size and how core the topic is to common customer needs — over strict recency, though most weeks the volume of flagged gaps is small enough that prioritization barely matters in practice.

Is the Content Gap Analyzer only useful for finding missing documentation, or can it also surface documentation that's actively wrong rather than just absent? It can surface both, though the clustering signal looks the same either way — low confidence — and distinguishing "missing" from "present but wrong" during review is what determines whether the fix is writing something new or correcting something that already exists.

How many low-confidence instances does it take before something gets flagged as a cluster worth reviewing? The threshold is tunable, but three or more instances of a similar low-confidence pattern within a week is the default we recommend, based on what's consistently distinguished a real pattern from noise across the deployments we've watched.

Is this something that only matters for large deployments with high ticket volume, or does it help smaller teams too? It helps smaller teams arguably more, in relative terms — a smaller team has fewer conversations to manually review for patterns, so a tool that automatically clusters and surfaces gaps replaces work that would otherwise require someone to notice a pattern across a much smaller, harder-to-spot sample.

The single habit underneath everything described in this piece, if we had to reduce it to one sentence, is this: treat every low-confidence cluster as a question worth a definite yes-or-no answer within the same week it's flagged, rather than letting it sit as an open item on a list that grows faster than it shrinks. Every team we've seen struggle with this tool wasn't struggling with the tool itself — they'd simply let the review cadence slip, and gaps that would have taken fifteen minutes to fix when fresh became harder to prioritize once a backlog of them had accumulated.

One pattern worth naming explicitly, since it's easy to miss until you've watched it happen a few times: the same underlying documentation gap sometimes surfaces as two or three seemingly unrelated clusters before anyone connects them, simply because customers phrase the same underlying confusion in genuinely different ways. A gap about a loyalty-discount stacking rule, for instance, might surface separately as a cluster about "can I use two codes" and a completely separately-worded cluster about "why didn't my discount apply," with nothing in the raw wording suggesting they share a root cause.

Catching that kind of connection is part of why we recommend a human doing the weekly review rather than fully automating the fix-decision step — an automated system clustering by wording alone would likely treat those two clusters as unrelated, while a person familiar with the actual policy can recognize the shared root cause immediately and fix both with the same single documentation change.

The habit is deceptively simple to describe and genuinely easy to let slip in practice, which is exactly why we keep returning to it throughout this piece rather than treating it as a settled, one-time recommendation: a fifteen-minute weekly review, done consistently, outperforms an occasional deep audit done sporadically, because gaps compound in cost the longer they sit undetected, in a way a sporadic audit, however thorough, can't fully make up for after the fact.

We'd add one final practical note: the report itself is only as useful as the review it triggers, which is why we've deliberately kept the weekly review lightweight enough that skipping it never has a good excuse — fifteen minutes, one person, a fixed time on the calendar, treated with the same seriousness as any other recurring operational commitment the team already honors without question.