All articlesInterviews

How three support leads run their AI Agents differently

Diego Ferreira10 min read
Interviews

Same product, three very different setups. What separated the team that trusted their Agent with refunds from the one that still reviews every reply turned out to have very little to do with risk tolerance as a personality trait, and much more to do with a specific, learnable difference in how each team had built its documentation.

Three approaches, side by side

AI would say"Refund approved"vsHuman did"Refund approved"

Three teams, same product, three completely different comfort levels. Team A lets the Agent issue refunds up to $200 with zero review. Team B requires sign-off on anything involving money. Team C hasn't turned on any write actions at all — the Agent only answers questions. On paper, that looks like three different philosophies about AI autonomy. In practice, talking to each team at length, it turned out to be something more specific and more fixable than a philosophical difference.

What actually separated them

The difference wasn't caution level, it was documentation maturity. Team A had a refund policy so explicit the Agent rarely hit an edge case; Team C's policy lived in three people's heads and a Slack channel. That's not a small distinction — it means Team C's actual constraint wasn't a lack of confidence in AI generally, it was a genuine absence of a written, unambiguous policy for the Agent to be confident about in the first place. No amount of AI capability closes that specific gap; only writing the policy down does.

Team B sat in between, and their specific hesitation was revealing: their policy was reasonably well documented, but specifically around whether money should ever move without a second set of eyes, independent of how confident the underlying eligibility check was. That's a legitimate, considered position rather than a documentation gap — some teams will reasonably keep a human in the loop on money movement even once everything else about the system has earned their trust, and Team B's setup reflects a deliberate choice rather than an unresolved uncertainty.

What all three agreed on

All three said the same thing about the first month: turn on fewer channels than you think you need, watch the trace on every automated action for the first two weeks, then expand once the pattern of mistakes — if any — becomes obvious. This consistency across three teams with otherwise very different setups is itself a useful signal — it suggests the "start narrow, watch closely, expand deliberately" pattern isn't specific to any one team's risk tolerance, it's closer to a universal best practice for onboarding an Agent regardless of how much autonomy you eventually plan to grant it.

The team running refunds unsupervised didn't start there. They spent three weeks in shadow mode, comparing what the Agent would have done against what a human actually did, before flipping the switch. Shadow mode specifically — running the Agent's proposed decision alongside the human's actual decision without acting on the Agent's output — gave them a concrete, measured comparison rather than a subjective sense of readiness, and it's the specific practice we now recommend to every team asking how to know when they're ready to expand autonomy.

What's notable in retrospect is that none of the three teams described their current setup as final. Team C, since this conversation happened, has started the same documentation-writing process that let Team A move faster, specifically because seeing the comparison made clear that their caution wasn't really a policy choice — it was a documentation gap they hadn't previously framed that way.

Common questions

How long does it typically take a team like Team C to reach documentation maturity comparable to Team A? It varies with how much undocumented tacit knowledge exists, but teams that treat it as a deliberate project — systematically interviewing the people who currently hold that knowledge in their heads and writing it down — tend to close most of the gap within a few weeks, rather than the months it might take if the writing happens only incidentally alongside other work.

Is shadow mode something every team should run, even ones that already feel confident in their documentation? We recommend it regardless of documentation confidence, because shadow mode also catches integration and edge-case issues that no amount of policy-writing surfaces in advance — it's a check on the system as built, not just the policy as written.

A longer look at Team B's specific reasoning

Team B — the one requiring sign-off on anything involving money regardless of confidence — is worth a longer look than the original comparison gave it, because their reasoning turned out to be more specific and more interesting than "we're cautious about money in general."

Talking to Team B's lead in more depth, the actual concern wasn't about the Agent's accuracy at all — by their own account, the accuracy on money-related decisions was already comparable to Team A's. The concern was about a specific category of edge case that had happened twice in their history before they'd even had an Agent: a customer whose account had been compromised, making a request that looked, on every available signal, like a legitimate refund request from the account owner, but wasn't.

That specific failure mode — a technically correct-looking request that's actually fraudulent, precisely because the fraud is good enough to pass every check a policy engine would apply — is one no amount of eligibility-checking sophistication fully closes, because the fraud is specifically designed to look eligible. Team B's human-sign-off requirement on money movement exists as a deliberate, permanent second layer against exactly that category of risk, not as a general lack of trust in the Agent's judgment for ordinary cases.

This distinction mattered enough that we've since started asking every team, during onboarding, to separate two different questions that often get conflated: "do you trust the Agent's judgment on ordinary eligibility calls" and "do you want a standing human check specifically against account-compromise-style fraud, independent of how good that judgment is." A team can answer yes to the first and still reasonably want the second, the way Team B does, and that's a coherent, defensible position rather than an unresolved trust issue.

Team A, when we described Team B's specific reasoning to them, said they'd made a different bet on the same risk: rather than a standing human sign-off on every transaction, they rely on a separate, automated fraud-signal check — unusual account activity, mismatched device fingerprints, that kind of thing — that runs alongside the ordinary eligibility check specifically to catch account-compromise patterns, escalating only when that specific signal fires rather than on every transaction regardless of signal.

Both approaches are defensible responses to the same underlying risk, and neither team considered the other's approach obviously wrong once the reasoning was laid out explicitly — which reinforced, for us, that "how much autonomy to grant an Agent handling money" isn't a single question with a single right answer, but several separable questions about which specific risks a team wants a standing human check against, each of which can reasonably be answered differently by different teams with different risk histories.

Common questions

Is there a risk that a team stays too conservative for too long, like Team C, simply out of inertia rather than a genuinely considered risk decision? This is common, and it's part of why we now ask teams directly and specifically what's actually driving their current level of caution — documentation gap, a specific risk like Team B's, or genuine inertia — rather than accepting "we're just cautious" as a sufficient answer on its own.

How do you help a team distinguish whether their hesitation is really about documentation, the way Team C's turned out to be, versus a genuine risk concern like Team B's? We walk through the same three-question framing described in this piece directly with the team — ask specifically what would need to be true for them to trust a given action type, and check whether the honest answer is "a documented policy exists" (a documentation gap) or "a specific risk we want a standing check against regardless of policy clarity" (a Team-B-style genuine concern).

Does a team's approach to autonomy typically stay consistent across all action types, or does it vary meaningfully by category even within one team? It varies meaningfully within a single team in most cases we've seen — a team might trust the Agent fully with order-status lookups while remaining Team-B-style conservative specifically on money movement, which is a coherent, common pattern rather than an inconsistency worth resolving.

How do you know when a team's caution has genuinely paid off versus just avoided a problem they'd never have hit anyway? This is hard to prove definitively in the absence of the counterfactual, which is part of why we lean on comparisons like Team B's specific account-compromise history — a team with a documented, specific past incident has real evidence their caution addresses a real risk, versus caution that's more speculative.

Is there a standard timeline for how long a team should expect to stay in a more conservative posture before reasonably considering expanding autonomy? There's no fixed timeline that applies universally — the shadow-mode comparison described in this piece is a better signal than elapsed time, since a team that's run enough shadow-mode volume to see a consistent pattern of agreement between the Agent's proposed decisions and actual human decisions has real evidence to act on, regardless of how many weeks that took to accumulate.

Does a team ever go the other direction — start with high autonomy and pull it back after a bad experience? It happens, and it's not a failure when it does — a team that grants broad autonomy, hits a specific bad case, and tightens the relevant policy in response is doing exactly the same kind of iterative calibration as a team that expands gradually, just discovering the boundary from the other direction.

What's the most common mistake teams make when trying to move from Team C's setup toward Team A's? Trying to write comprehensive documentation for every possible case before turning anything on, rather than starting with the most common, best-understood cases and expanding coverage incrementally as shadow-mode data reveals what's actually missing.

If there's one thing we'd want a support lead reading this to take away, it's to stop asking "how much do we trust AI in general" as a single question, and start asking the more specific version for each action type separately: is our hesitation here about documentation, about a specific historical risk like Team B's, or genuinely just inertia. Three very different setups turned out to make sense for three teams once each answered that more specific question honestly for themselves, rather than defaulting to a generic comfort level applied uniformly across every kind of decision the Agent might make.

We'd add one more observation from following up with these three teams roughly a year after the original conversation: none of the three setups described above turned out to be permanent. Team A's has stayed roughly the same, since their documentation-driven confidence has continued to hold up. Team B has since automated a narrower category of low-dollar refunds specifically, while keeping their sign-off requirement for everything above that threshold — a partial evolution rather than an all-or-nothing shift. And Team C, as mentioned, is now roughly where Team A was a year ago, having closed most of the documentation gap that was the real constraint the whole time.

What we'd want any support lead reading this to sit with isn't which of the three teams they most resemble today, but which specific question — documentation gap, a real historical risk, or plain inertia — actually explains their own current hesitation, because that answer, more than any comparison to these three teams specifically, is what determines what their own next reasonable step should be.

We'll close by noting that none of the three teams described here would claim to have the definitively "correct" setup — each considers their own current approach the right fit for their own specific documentation maturity and risk history, which is exactly the point this piece has tried to make: there isn't a universal right answer to how much autonomy to grant, only a right answer for a specific team's specific, honestly-assessed situation.