Multi-agent isn't a personality split — and context isn't free
"Multi-agent" has started to sound like the new "microservices" — sometimes exactly the right call, often just one agent's job split into four pieces for a slide. We went looking for a way to tell the difference before building anything, and ended up reaching for an unlikely source: the segmentation, targeting, and positioning exercise marketers use to decide who to talk to and how. It mapped surprisingly well onto deciding how to split up an agent's responsibilities — and onto a second, related question we kept running into: how much context each piece should actually see.
We're borrowing the framework loosely, not literally — there's no actual customer segment here, only decisions about which piece of a system should own which kind of judgment. But the discipline of forcing three separate, explicit decisions instead of one vague one turned out to transfer cleanly.
Segmentation
The first useful move was refusing to split by persona and instead splitting by concern. Four personas answering the same underlying question isn't a segmentation, it's a costume change, and we caught ourselves doing exactly that in an early prototype — four differently-named "agents" that all had access to the same tools and the same context, distinguished only by a different system prompt describing a different tone. Nothing about that setup actually reduced risk or improved a decision; it just added four times the prompts to maintain.
The segmentation that actually held up looked like this: one piece classifies intent, one owns policy and eligibility, one owns tool execution, and one owns customer-facing tone. Each of those is a genuinely different job with a different failure mode — a tone mistake is embarrassing, a policy mistake is expensive, a tool-execution mistake can be irreversible — and that's the test we now apply before splitting anything: does this piece fail differently than the others, or is it just wearing a different name for the same underlying judgment.
The test also works in reverse, as a reason not to split something further. We considered separating "policy and eligibility" into two pieces — one for eligibility checks, one for dollar-limit checks — and decided against it, because both fail the same way (an incorrect authorization) and both draw on the same underlying data. Splitting them would have added coordination overhead without adding a genuinely different failure mode to isolate.
Targeting
Once the segments existed, the harder question was deciding what each one should actually be responsible for deciding — targeting the right scope of judgment to the right piece. The clearest case for this showed up in ambiguous tickets. "This isn't what I ordered" could mean the wrong item, the wrong size, a mismatch with expectations, or plain buyer's remorse — four different resolutions hiding behind four identical words, and no amount of intent classification alone resolves that ambiguity, because the ambiguity is real, not a parsing failure that better prompting fixes.
We targeted that specific decision — should the agent answer now or ask first — at a dedicated confidence check rather than leaving it implicit in the main response. Below a threshold, it asks one clarifying question instead of guessing; resolution on genuinely ambiguous tickets went up once that check existed as its own step instead of being folded into "just answer well." One good question, targeted at the right moment by the piece responsible for that decision, beats a confident wrong answer almost every time.
The targeting question also applies to which piece decides when to stop trying and hand off. We initially let the tone-owning piece make that call, on the theory that it was closest to reading the customer's frustration level. It made the call badly, because it had no visibility into whether the underlying issue was actually resolvable — it was targeting a decision it didn't have the right information to make. Moving that decision to the policy piece, which does have that visibility, fixed it without changing a single line of the tone logic.
Positioning
The last piece was positioning — deciding how much of the world each piece should see, since visibility is itself a kind of framing that shapes what a piece concludes. We tested giving the agent a customer's full purchase history against giving it only the order relevant to the current ticket. Full history occasionally surfaced an unrelated past complaint and visibly derailed the conversation — in one transcript, the agent referenced a shipping delay from three months earlier that the customer had already forgotten about, and the customer's reply made it clear the reference read as odd rather than helpful. The scoped version stayed on topic and resolved faster.
Position a piece with too much context and it starts answering the wrong question well instead of the right question at all — bigger context windows don't make an agent smarter by default, they make it easier to wander into territory nobody asked about. We now default every piece to the narrowest context that lets it do its specific job, and treat any request to widen that scope as something that needs its own justification, not a default we grant because the window technically allows it.
The same positioning problem shows up even earlier, in ordinary memory rather than tool design. A customer repeating themselves to a bot isn't a UX bug, it's a memory architecture bug — if the agent forgets what happened two messages ago, no amount of prompt tuning downstream fixes that upstream gap. Segmenting concerns, targeting the right decision to the right piece, and positioning each piece with the right amount of context turned out to be one discipline wearing three names, applied at three different points in the same pipeline.
In practice
A concrete split
Here's what this looked like on one real ticket type — a return request where the customer wants a different size instead of a refund. The intent-classifying piece recognized this as an exchange request, not a refund request, from the first message; getting that classification wrong at the start would have sent the whole conversation down the wrong path regardless of how well anything downstream performed.
The policy-and-eligibility piece then checked whether the specific item was eligible for exchange — some categories aren't, based on return-window and item-condition rules that have nothing to do with language understanding and everything to do with business rules that change independently of any model. This is also where the piece checked stock on the requested replacement size, a lookup that has nothing to do with the customer's intent and everything to do with a live inventory system.
The tool-execution piece, once eligibility cleared, made the actual calls: reserve the replacement item, generate a return label for the original, and update the order record — three separate API calls, each of which could independently fail, and each of which the execution piece was responsible for retrying or rolling back correctly, a job that has nothing to do with understanding language and everything to do with handling distributed-system failures gracefully.
The tone-owning piece took the outcome of all of that — success, partial failure, or a stock-out that meant the exchange couldn't happen as requested — and turned it into an actual message to the customer, including the harder case of explaining a stock-out without it reading as a form rejection. That piece never touched the inventory API and never decided eligibility; its only job was turning a structured outcome into language a person would actually want to read.
Watching this split hold up across a few thousand real exchange requests is what convinced us the segmentation was right, more than any argument we could have made about it in the abstract. Each piece failed independently exactly as predicted — a stock-lookup timeout never corrupted the tone of the final message, and a genuinely confusing customer message never caused an unauthorized inventory reservation.
Common questions
How do you know when four pieces is too many, versus not enough? We watch for coordination cost specifically — how often two pieces need to pass information back and forth more than once to resolve a single ticket. A jump in that number is usually a sign two pieces should be merged, because they're functionally deciding the same thing from two different angles instead of owning genuinely separate decisions.
Doesn't scoping context this tightly risk the agent missing something important from the customer's history? It's a real trade, and we made it deliberately in the direction of missing an occasional relevant detail rather than frequently surfacing an irrelevant one. The cases where full history would have genuinely helped were rare in our data; the cases where it actively confused the conversation were common enough to notice within the first few hundred tickets we reviewed.
Is this approach specific to support, or does it generalize to other kinds of agents? The specific split we landed on is specific to support, but the three questions behind it aren't: does this fail differently, does this piece have the right information to decide, and what's the narrowest context it needs. We'd expect those same three questions to produce a different, equally valid split for a different domain, which is exactly why we treat the framework as reusable and the specific segmentation as not.
What's the failure mode you'd most want a team to avoid when trying this themselves? Splitting by persona instead of by concern, which we did ourselves in the first prototype. It feels like progress because you now have four named things instead of one, but if all four can see the same tools and the same context and differ only in tone, you've added maintenance burden without adding any actual safety or clarity.
What changed after we shipped this
The segmentation held up well enough in production that the more interesting story became what happened when we tried to add a fifth concern later — refund-amount escalation, for cases above a certain dollar threshold that needed a different kind of review than ordinary policy checks.
Our first instinct was to add this as a rule inside the existing policy-and-eligibility piece, since it's clearly a policy-adjacent decision. It technically worked, but it made that piece's failure mode less clean: a bug in the escalation-threshold logic now showed up as a policy failure indistinguishable, in our monitoring, from an ordinary eligibility failure, even though the two have almost nothing in common operationally — one is about whether an action is allowed at all, the other is about who gets to approve an allowed action.
We eventually split it out as its own piece, applying the same segmentation test we described earlier: does it fail differently (yes — an escalation-threshold bug produces a wrongly-skipped human review, not a wrongly-approved refund), does it have the right information to decide (yes, once given the proposed action and its dollar amount), and what's the narrowest context it needs (just the action and the amount, nothing about the conversation's tone or the customer's history). All three answers pointed the same direction, which is roughly the confirmation we look for now before deciding whether something is really a new concern or just a rule that belongs inside an existing one.
That experience changed how we onboard a new concern generally. We no longer ask "does this feel like it belongs with an existing piece" as the first question, because that instinct is exactly what led us to bolt escalation logic onto the policy piece in the first place. We ask the three-question test explicitly, every time, even when the answer feels obvious going in — it wasn't obvious to us the first time, and the cost of getting it wrong compounds the longer a poorly-placed rule sits inside the wrong piece.
The broader lesson, if you're designing something similar: expect to get the initial split wrong at least once as the system grows past its first few concerns, and build in the expectation that splitting a piece further, later, should be a normal, low-drama operation rather than something that requires re-architecting. The pieces that were easiest to split later were the ones we'd kept smallest and most single-purpose from the start, even when a slightly larger piece would have felt more efficient to build initially.
Before splitting anything into a separate piece now, we ask three questions in order: does it fail differently than the neighboring pieces (segmentation), is it the piece with the actual information needed to make this specific call (targeting), and what's the narrowest context it can do its job with (positioning). A piece that fails the first question gets merged back in. A piece that fails the second gets its decision moved somewhere else. A piece that fails the third gets its context trimmed until it stops wandering.
None of this is really about "multi-agent" as a category. It's about refusing to let architecture decisions default to whatever's easiest to prototype — four chatty personas — instead of what's actually structurally different underneath.