What actually breaks chatbots: a conversation with our head of support
We asked the person who reviews the most AI conversations on the platform what still trips agents up, expecting a fairly technical answer about model limitations. What we got instead was a list of conversation-design problems that have almost nothing to do with the underlying model's raw capability.
The interview below is lightly edited for length, but the substance and ordering of what came up first is unchanged from how the conversation actually went.
On the biggest recurring problem
"Multi-part questions are the number one thing that trips an Agent up — someone asking about a refund and a shipping delay and a coupon code in the same message. The Agent has to correctly split that into three separate things to check, and if it collapses them into one, it usually ends up half-answering all three instead of fully answering any of them. What's interesting is that this isn't really a language-understanding failure — the Agent usually does parse out that there are three separate asks. The failure is more often in how the response gets structured afterward, addressing them out of order or blending them together in a way that reads as confusing even when every individual fact stated is correct."
On silence and re-engagement
"The second is silence. A customer who doesn't respond to a clarifying question for two days, then replies as if no time passed. Agents that don't re-establish context before continuing give answers that make no sense, because they're answering the two-day-old question in isolation rather than acknowledging that time has passed and possibly checking whether the original context is even still relevant. A customer coming back after two days with 'yes' to a question the Agent asked two days ago deserves a brief re-grounding — 'just to confirm, you're saying yes to the replacement option for the blue jacket, order 48213?' — not a bare continuation as if the gap never happened."
On adversarial testing
"The third, and the one that surprised me most, is customers testing the Agent on purpose — asking if it's a robot, trying to get it to say something off-brand. Those conversations need a different tone entirely than a real support question. Early on, we saw Agents respond to these the same way they'd respond to a genuine question, which reads as either evasive or oddly literal, and neither lands well. The fix wasn't a smarter model — it was recognizing this as its own conversational category and giving it its own handling, roughly: acknowledge directly, don't get defensive, and redirect toward something the Agent can actually help with."
On what this means for building agents
None of these are model problems. They're conversation-design problems, and fixing them looks more like writing better prompts and handoff rules than retraining anything — which is a useful thing to hear from someone who reviews thousands of real conversations, because it's easy for teams building agents to assume every failure they see is a capability ceiling, when a meaningful share of what actually breaks is closer to a UX problem wearing an AI costume.
What ties the three failure modes together, on reflection, is that all three involve the Agent losing track of something that isn't in the most recent message — multiple intents bundled together, elapsed time since the last exchange, or the meta-context that a message is testing the system rather than genuinely asking something. Each of those requires the Agent to reason about the conversation as a whole, not just the current turn, which is exactly the kind of state-tracking discipline we've written about elsewhere in this series applying to ticket resolution specifically.
Common questions
How do you train an Agent to recognize when it's being deliberately tested rather than asked a genuine question? It's less about training a classifier for "is this a test" and more about pattern-matching a small set of recognizable signals — direct questions about whether it's a robot, requests to say something clearly off-brand — and having a specific, calm response ready for that category rather than routing it through the same reasoning as an ordinary support question.
Does the multi-part-question problem get worse as conversations get longer and more complex? It compounds, yes — a conversation that's already juggling two unresolved threads is more likely to mishandle a third one introduced mid-conversation than a fresh message would be, which is part of why explicit conversation state, tracked separately from the raw message history, matters as much as it does.
What changed once the team started tracking these three patterns explicitly
Once the head of support started reviewing transcripts specifically for these three failure categories — bundled multi-part questions, silence-then-return gaps, and adversarial testing — rather than for general quality, a fourth, related pattern emerged that hadn't been named going into the conversation: conversations where the customer switched topics entirely mid-thread, without any elapsed time gap, simply because a second, unrelated issue occurred to them while the first was still being resolved.
That fourth pattern shares a root cause with the bundled-question problem — both involve more than one distinct thread of intent inside a single conversation — but it's structurally different in a way that mattered for the fix. A bundled question arrives as one message with multiple parts already inside it. A mid-thread topic switch arrives as a new message that introduces something entirely separate from what the conversation had been about up to that point, and the Agent has to recognize that shift rather than trying to fold the new topic into the still-unresolved first one.
The fix that worked for bundled questions — explicitly parsing out separate intents and tracking them individually — needed a small extension to handle mid-thread switches gracefully: rather than assuming every new message is either a continuation or a fully separate conversation, the Agent now explicitly checks whether a new message plausibly continues the current tracked intent or introduces a new one, and if it's ambiguous, asks a brief clarifying question rather than guessing.
Watching this fix land in practice, the head of support specifically called out one transcript where a customer asked about a shipping delay, then, before that was resolved, added "also my discount code isn't working" in a follow-up message. The updated handling recognized this as a second, separate intent rather than trying to merge it into the shipping conversation, tracked both threads independently, and resolved the discount-code issue first, since it happened to be quicker, before returning to finish the shipping question — closing both cleanly rather than producing a single muddled response trying to address both at once.
The adversarial-testing category also evolved with more data. Beyond the direct "are you a robot" style test, a subtler version showed up repeatedly: customers deliberately feeding the Agent contradictory information to see whether it would notice — claiming a different order number than the one on file, for instance, specifically to probe whether verification was actually being enforced or just assumed. Treating this as its own recognizable pattern, distinct from an honest customer simply misremembering a number, required building a specific response for a mismatch: verify calmly, ask for confirmation, and if the mismatch persists, treat it as a potential account-security signal rather than either accusing the customer or quietly working around the discrepancy.
None of these refinements required a different underlying model. Every one of them came from the same discipline the original interview pointed at: watching real conversations specifically for known failure categories, naming new ones as they emerge, and treating conversation design — not model capability — as the lever that actually moves the needle on most of what breaks a chatbot in practice.
Common questions
How do you train a support team to recognize which of these failure categories a bad transcript actually falls into, rather than just noting "the Agent handled this poorly"? We built a specific tagging taxonomy around the categories described in this piece — bundled intent, elapsed-time gap, adversarial test, mid-thread topic switch — and require reviewers to tag transcripts against it specifically, rather than leaving quality assessment as a vague, unstructured judgment.
Do these same failure categories apply equally across text-based chat and voice-based agents, or does voice introduce its own distinct patterns? The core categories transfer, but voice adds its own wrinkle around disfluency and interruption that text doesn't have in the same way — a customer talking over the Agent mid-response is a distinct failure mode from any of the four described here, and it's on our list to study with the same rigor once we have enough voice-specific transcript volume.
Is there a risk that explicitly designing for adversarial testing makes the Agent feel evasive or overly guarded even to customers who aren't testing it? This is a real design tension, and the fix is keeping the adversarial-response handling narrow and specific to clearly recognizable test patterns, rather than making the Agent broadly suspicious or hedgy in ordinary conversation — a false positive here, treating a genuine question as a test, is its own kind of failure worth watching for.
How often do these four failure categories actually occur relative to conversations that go smoothly? They're a minority of overall conversations, but a disproportionate share of the ones that end up escalated or complained about, which is exactly why they're worth this much specific attention despite being relatively rare in raw frequency terms.
Does fixing one of these categories ever make another one worse, or are the fixes independent of each other? Mostly independent in our experience, though the mid-thread topic-switch handling and the bundled-intent handling share enough underlying logic — both about tracking multiple simultaneous intents — that a fix to one occasionally requires re-testing the other to make sure the shared logic still handles both cases correctly.
Is silence-then-return a bigger problem in some channels than others? Email and ticket-based channels see it far more than live chat, simply because the medium itself invites longer gaps between messages — chat conversations that go silent for two days are usually just abandoned, while an email thread picking back up after two days is completely normal user behavior that the Agent needs to handle gracefully.
What's the single most useful thing a team can do to reduce these specific failure modes? Review real transcripts specifically looking for these three patterns, rather than reviewing for general quality — a targeted review looking for "did this collapse multiple intents," "did this handle an elapsed-time gap," and "did this handle an adversarial test" surfaces very different, more actionable findings than a general read-through does.
Looking back at all four failure categories together — bundled intent, elapsed-time gaps, adversarial testing, and mid-thread topic switches — the thread connecting them is that every one involves the Agent needing to reason about the conversation as a structured whole rather than as a sequence of independent messages. That's a more demanding requirement than it sounds like, and it's the reason none of these problems got solved by a better underlying model on its own — they got solved by explicitly building conversation structure into the system, the same structural discipline we've described in more depth elsewhere in this series applied to ticket resolution specifically.
We'd expect any team building a support Agent to eventually run into some version of all four categories described here, regardless of which underlying model they're using, precisely because these are properties of multi-turn conversation itself rather than properties of any particular model's limitations.
It's worth adding one caveat to how we described the fix for each category above: none of these fixes were a single change deployed once and considered done. Each of the four failure categories has gone through at least two or three rounds of refinement as new specific instances surfaced edge cases the first fix hadn't anticipated — the mid-thread topic-switch handling described in this piece, for instance, needed a follow-up adjustment for cases involving three or more simultaneous threads rather than the two-thread case it was originally built and tested against.
That iterative pattern is, in our experience, normal and expected for this kind of conversation-design work generally — the first fix for a named failure category closes the specific instance that revealed the category, and later instances reveal edge cases within that category that need their own follow-up refinement, in a way that never fully terminates but does produce diminishing returns of new edge cases over time.