All articlesEngineering

Anatomy of three support tickets that looked simple and weren't

Marcus Ibe11 min read
Engineering

"I was charged twice for the same order." Nine words. If you've never had to build the system that resolves that sentence correctly, it's easy to assume there's not much to it — read the ticket, look up the order, refund the extra charge, done. There is a lot to it, and walking through what actually happens underneath three ordinary-looking tickets is the fastest way we've found to explain why support automation is a systems problem, not a chat problem, structured the way a lot of persuasion is structured: attention, interest, desire, action.

We picked these three because none of them look hard from the outside, and all three route through a different combination of systems once you actually build the resolution path properly instead of demoing the happy case.

Attention

Start with the duplicate charge, because it's the one that looks simplest and still isn't. A customer sends nine words. Behind those nine words, resolving the ticket correctly takes five distinct steps: pull both charges and confirm they share an order id and amount but different timestamps, check policy for whether a duplicate within 24 hours auto-refunds or needs review, issue the refund through the payment API rather than promising one in the chat window, confirm the refund status actually came back as succeeded — not just that the API call returned without an error — and only then tell the customer, with the real amount and the real timeline, not a placeholder sentence that sounds reassuring.

Skip the confirmation step specifically and you've told someone their money is coming back when the underlying call silently failed partway through — a timeout on the payment processor's side, a rate limit, a stale idempotency key. We found this exact failure mode in testing by deliberately injecting a slow response from a sandboxed payment API and watching what the agent told the customer before we'd built the verification step. It told them the refund was on its way. It wasn't.

That's the detail that earns this ticket its place first: the version of "duplicate charge" that looks trivial in a product demo is the version where every step happens instantly and successfully. The version that actually determines whether customers trust the system is the version where step four takes eight seconds and the agent has to decide what to say in the meantime — and the honest answer, "still confirming, one moment," turned out to build more trust than a premature "all set!" that occasionally wasn't true.

Interest

Now take a ticket that reads almost identically on the surface but resolves nothing like it. "My package says delivered but I never got it" starts as a fact-finding question, not a refund-eligibility one, and treating it as the latter from the first message is the single most common mistake we saw in earlier versions of this flow. The correct first move is to pull carrier tracking and any delivery photo or signature on file, then check whether that shipping address has a history of missing-package claims — a pattern, not proof of anything on its own, but a signal worth weighing before deciding what happens next.

With no pattern, policy allows an automatic reship or refund with no human involved, because the evidence on file supports the customer's account and nothing about the request looks anomalous. With a pattern — say, a third missing-package claim at the same address inside two months — the case goes to a human with that history attached, and nothing gets auto-approved, because now the same nine words carry a materially different kind of risk than they did the first time someone said them.

Two tickets, nearly the same wording, completely different resolution paths — decided entirely by data the customer never sees and would have no reason to mention themselves even if they wanted to. That's the part worth genuine interest: the words in the ticket are a weak signal compared to the context sitting behind them, and a system that only reads the words is guessing at a problem that's already been answered by data it hasn't bothered to check.

We also learned, the hard way, that the order these checks run in matters. An earlier version pulled the address history before confirming the tracking status, which meant a customer with a clean history but a genuinely lost package sometimes got auto-approved before the agent had even confirmed the package showed as delivered in the first place — technically correct in outcome, but built on the wrong evidence, and one bad carrier data sync away from being wrong in outcome too.

Desire

Here's why this is worth caring about even if you're not the one building it. Neither of the two resolutions above works without a conversation actually having a state — open, awaiting customer, awaiting system, resolved, escalated — a real position in a real process, not just a running list of messages accumulating in a thread. An agent that only tracks "messages so far" doesn't know where it is; it's reacting to whatever was said last instead of knowing what step comes next, and that's exactly where a five-step resolution collapses back into three confused, out-of-order ones.

We saw this concretely in an early version of the flow: a customer asked a follow-up question mid-resolution — "how long will the refund take?" — and the agent, having no state to fall back on, restarted its reasoning from the follow-up question alone, momentarily forgetting it had already verified the duplicate charge and was mid-refund. It answered the timing question fine. It also, in the same reply, asked the customer to confirm the order number again, as if the conversation were starting over. The fix wasn't a smarter prompt. It was giving the conversation an explicit state that survives being interrupted by an unrelated question.

The same discipline applies when a ticket does need a human, and this is the part that determines whether "resolved by AI" numbers mean anything real. We treat escalation as a handoff, not a dead end — the case arrives with what the customer asked, what the agent already checked, what it tried, and why it stopped short. A human picking up a case cold, with none of that context, ends up re-doing the fact-finding the agent already completed, which is slower for the customer than if the agent had never touched the case at all. That's the agent's failure to design a proper handoff, not the human's failure to be efficient, and it's the gap most "AI resolved X% of tickets" numbers quietly paper over.

Action

What ties all three tickets together isn't the individual steps — it's how consistently a single customer sentence turns out to route through three or four systems before anything is actually resolved: a tracking or payment API, a fraud or pattern signal, a policy engine, a refund or reship flow, sometimes an escalation queue at the end holding the case together with the context it's accumulated along the way.

If you're building or buying agent automation, that's the concrete question worth asking about any specific ticket type before trusting a resolution-rate number attached to it: not "can it answer this," but "how many of those systems does answering this correctly actually require, and does the agent's resolution path touch all of them or just the first one it happens to reach." A demo that only shows the happy path for a duplicate charge — order found, refund issued, done — has shown you roughly a fifth of what the real ticket requires.

In practice

The pattern we now use when scoping any new ticket type before building the automated path for it: write out every branch a human agent would actually consider, not just the branch that ends in the obvious action. "Package says delivered, customer says otherwise" has at minimum two branches — no fraud signal, fraud signal — and each branch needs its own explicit policy, not a single prompt asked to weigh both possibilities silently and pick one on its own.

Common questions

Doesn't this level of detail only matter at large scale? It matters the first time any single one of these tickets goes wrong, which can happen on ticket number twelve just as easily as ticket number twelve thousand. Scale changes how often you notice a gap like this, not whether the gap exists. A five-step resolution with one step missing fails at any volume; it just fails more visibly, and more often, once volume is high enough that the missing step gets tested by real edge cases instead of the happy path you built and tested against.

How do you decide how many branches to build for a new ticket type, instead of trying to cover every possible case? We build for the branches that change the resolution path, not every branch that changes the wording. "My package says delivered but I never got it" and "my package never arrived" are different words for the same branch if the resolution is identical either way. The fraud-signal branch and the no-signal branch are worth building separately because the resolution genuinely differs. The test isn't how many ways a customer might phrase something — it's how many different correct actions exist underneath those phrasings.

What happens when a ticket doesn't fit any of the branches you've built? It should fail toward a human, loudly, rather than toward the nearest branch that looks close enough. We deliberately made the classification step conservative about what counts as a match for an existing branch, because a near-match treated as an exact match is how a five-step resolution silently turns into the wrong five steps, executed confidently.

Is state tracking worth building before you have any real traffic to learn from? Yes, more so than almost anything else on this list. We built it after the fact specifically because we didn't have it from the start, and the retrofit cost more engineering time than every other piece of this flow combined. State tracking doesn't need real traffic to design correctly — it needs someone to write out what states a conversation of this type actually passes through, which is a design exercise you can do on day one.

What we measure now

After building out all three ticket types properly, we changed which numbers we watch weekly, because the ones we started with — messages per ticket, time to first response — didn't actually tell us whether a resolution had gone through the right steps or just arrived at a plausible-looking outcome.

The number we added first was steps-skipped rate: how often a resolution completed without hitting every step in its defined path, caught by instrumenting each step explicitly rather than inferring it from the final outcome. A refund that goes through without the verification step still counts as "resolved" by a naive resolution-rate metric. It shows up immediately in steps-skipped rate, which is exactly the number that would have caught our early duplicate-charge failure before a customer did.

The second was branch-coverage drift: what percentage of a given ticket type's real traffic actually matches one of the branches we've explicitly built, versus falling into a generic catch-all. A rising catch-all percentage on any ticket type is an early warning that customer behavior has shifted in a way our branches haven't kept up with — new fraud patterns, a new product line with different return rules, a seasonal spike in a specific complaint we hadn't scoped for.

The third, and the one that took longest to justify building, was state-recovery accuracy: when a conversation gets interrupted by an unrelated question and then resumes, does the agent correctly return to the state it was in before the interruption, or does it quietly restart? We only started tracking this after finding the exact failure by accident during a transcript review, and it's now one of the numbers we watch most closely, because a low score there predicts confusing, repetitive customer experiences well before those show up in a satisfaction survey.

None of these three replaced resolution rate — they sit alongside it, specifically because resolution rate alone would have told us everything was fine in every case where it turned out not to be. The pattern we'd recommend to any team building something similar: for every ticket type, ask what a naive resolution-rate metric would fail to catch, and instrument specifically for that gap rather than assuming a single top-line number is enough.

The second habit is treating conversation state as a first-class object from day one, even for ticket types that look linear. It's cheap to add before a flow exists and expensive to retrofit once a few thousand conversations have already happened without it — we know because we did the retrofit, and it took longer than building the original flow, mostly because we had to reconstruct state retroactively from message history instead of just reading it off directly.

The chat bubble is the entrance. The actual product is everything that happens in the two seconds after — the systems it calls, the state it tracks, the branch it silently decides not to take — and that's true whether the ticket takes five steps or fifteen. Design for the fifteen-step version even when most tickets only need five; the five-step ones will still work, and the fifteen-step ones won't quietly fail in front of a customer.