All articlesInsights

Pricing an AI support agent: the questions that actually matter

Marcus Ibe10 min read
Insights

Pricing AI agents per message measures the wrong unit entirely. Nobody is buying messages — they're buying problems that stop existing — and a pricing model built around message volume quietly rewards the agent for talking more, not resolving more. Working through this properly meant treating the agent like an actual product with a price, a place it gets bought, and a promotion built around honest claims instead of vague ones.

We came to this the hard way, after watching our own internal usage dashboards for a few months and noticing that the metric climbing fastest — messages per resolved ticket — was one nobody had actually set out to optimize, and one that would have looked like growth on a per-message pricing model while actually representing the opposite of progress.

Product

The product being sold isn't "an AI that answers questions." It's a system that resolves a specific volume of specific problems, and the metric that should define the product is cost-to-serve sitting right next to resolution rate, not behind it. An agent that resolves 95% of tickets by calling twelve tools and retrying three times per conversation can end up costing more than the human it replaced — a high resolution number alone doesn't tell a buyer whether the underlying economics actually work, and "the product resolves things" isn't a complete claim without "at what cost per resolution."

We started reporting cost-to-serve internally as a per-ticket-type number rather than a single blended average, because the blended number hid something important: duplicate-charge tickets cost a fraction of account-takeover tickets to resolve correctly, and averaging them together made it look like the product had one cost profile when it actually had several, each shaped by how many systems that ticket type has to touch.

Price

Getting the price right meant testing an assumption rather than keeping it. We ran the same set of tickets through three routing strategies — always the strong model, always the cheap model, and routing by complexity — and measured cost per resolved ticket, not cost per call. Routing by complexity won on cost without giving up resolution rate. That result reframes which model actually wins on price: if a stronger model costs twice as much per call but resolves an issue without a human escalation attached, the total cost can land lower than the cheaper model that quietly pushes more conversations to a person. "Which model is cheapest" and "which model is cheapest per resolved problem" are different questions with different right answers, and only one of them should set the price.

This also changed how we think about discounting. A per-message price makes a volume discount look generous while actually rewarding whatever's driving message count up, including inefficiency. A per-resolution price makes a volume discount mean something real: a customer sending us more resolved problems is genuinely cheaper for us to serve at scale, and the discount reflects an actual cost curve instead of a vanity metric.

Place

Where this gets sold matters as much as what it costs. The conversation that actually lands with a support leader isn't "how many tickets can AI handle" — it's how many of their best agents are stuck doing L1 work they've already outgrown, and what a policy engine, escalation design, and honest cost-per-resolution reporting buy them once that work moves off their plate. That's a different sale, aimed at a different moment in a support leader's week, than a generic "automate your tickets" pitch, and it changes where and how the product should actually be positioned to be bought.

It also changes who's in the room for that conversation. A per-message pricing pitch tends to get evaluated by whoever owns the support budget line item. A cost-per-resolution pitch, once the buyer actually believes the number, tends to pull in whoever owns the broader efficiency or headcount conversation — a different, usually more senior, buyer with a different set of questions, mostly about the failure cases rather than the happy path.

Promotion

The promotion that holds up is the one that survives a hard question, because support leaders are going to ask it eventually. "What's your model" is the wrong first question for a buyer to ask a company building this, and it's the wrong first claim for that company to lead with. "What's your failure rate on the 10% of tickets that aren't the happy path" is the harder, better question — and the honest answer to it, published rather than hidden, is worth more as promotion than any message-volume pricing page, because it's the one claim a buyer can actually verify before they commit to a price.

We rewrote our own pricing page around this after realizing our first version led with the same message-volume framing we'd just spent months arguing against internally. The updated version leads with cost per resolved ticket by category, with the caveats attached rather than buried, and it's converted better with exactly the buyers we most wanted to reach — the ones who were going to ask the hard question anyway and would rather see it answered upfront.

Working through a real example

Take a mid-sized support operation handling roughly four thousand tickets a month, split across the ticket types we've discussed elsewhere in this series — duplicate charges, missing packages, account questions, exchanges. Priced per message, the highest-cost tickets to serve and the cheapest ones to serve look identical on the invoice, because the invoice only counts messages exchanged, not systems touched.

Priced per resolution, with cost-to-serve broken out by ticket type, the same four thousand tickets tell a completely different story: the duplicate-charge tickets, resolvable in a handful of steps through one or two systems, cost a fraction of the account-takeover tickets, which route through identity verification, security logging, and often a human, regardless of how efficiently the agent handles its part.

That breakdown changes the conversation with a buyer in a specific, useful way. Instead of negotiating a single blended rate, the conversation becomes about which ticket types the buyer most wants automated first — usually the cheap, high-volume ones — with an honest expectation about which ones will always cost more to resolve well, no matter how good the underlying model gets, because the cost lives in the number of systems touched, not in model quality.

We've found this breakdown also surfaces a useful signal for the buyer that a blended number hides: if account-takeover tickets are a large share of their volume, that's not primarily an AI problem to solve at all — it's a signal about something upstream, like weak authentication, that no amount of agent automation on the support side actually fixes. A per-resolution pricing conversation tends to surface that observation naturally. A per-message conversation almost never does.

The same real example is also where routing by complexity earns its keep most visibly. Running all four thousand tickets through the strongest available model regardless of type would resolve slightly more of the ambiguous ones, at a cost the volume of simple, cheap-to-resolve tickets doesn't justify paying for. Routing by complexity means the duplicate-charge tickets — the majority, in most operations we've looked at — get handled by whichever model is cheapest at that task without sacrificing anything, while the harder, rarer tickets get the more expensive model exactly where it earns its cost back in avoided escalations.

Common questions

Won't buyers just push back on paying more for the harder ticket types? Some will, and that's a useful filter. A buyer who wants a single flat rate regardless of ticket difficulty is usually a buyer who hasn't yet had to explain to their own leadership why the blended number moved when their ticket mix shifted — which happens constantly in a real support operation. The buyers who've already been burned by that conversation tend to prefer the breakdown, once they see it.

How do you avoid the pricing model becoming so granular it's confusing to buy? We cap the number of ticket-type tiers at what a buyer can actually reason about without a spreadsheet — a handful of categories, not dozens. The goal is honesty about cost variance, not maximum precision; a pricing model nobody can hold in their head defeats the purpose even if it's technically more accurate.

Does routing by complexity risk under-serving the hard tickets to save money? Only if complexity routing is built to optimize cost without a resolution-rate floor attached, which is why we test any routing change against both numbers together, never cost alone. A routing strategy that saves money by quietly resolving fewer hard tickets isn't cheaper — it's just moved the cost to an escalation queue and made it less visible.

Is there a risk that publishing failure rates by ticket type just hands ammunition to competitors? Less than it looks like there would be, because the number that matters isn't the raw failure rate — it's the trend, and what's changed because of it. A competitor can see a number; they can't as easily see the specific fixes behind why that number is lower than it was two quarters ago, which is the part that actually took the work to build.

The number that actually convinced our own team

Every argument in this piece was easier to make to customers than it was to make internally, at first, because our own sales team had gotten comfortable with a per-message pricing conversation that was simple to explain even though we'd already started to suspect it measured the wrong thing.

What actually shifted the internal conversation wasn't an argument at all — it was a single chart, built from our own usage data, plotting messages-per-resolution against ticket complexity across a full quarter. The line wasn't flat, which we expected. It was upward-sloping in a specific, uncomfortable way: the hardest tickets weren't just taking more messages because they were inherently harder, they were taking disproportionately more messages relative to how much harder they actually were, which meant something in the agent's approach to complex tickets was inefficient in a way a simple message-volume price would have quietly rewarded rather than caught.

That chart is what got budget approved to build the cost-to-serve breakdown by ticket type described earlier in this piece, because it made the abstract argument — "message volume isn't the right unit" — into a concrete, uncomfortable number that was actively growing every quarter we didn't address it. Abstract arguments about pricing philosophy move slowly inside a company that already has a working, if imperfect, pricing model. A chart showing the imperfection actively getting worse moves much faster.

The same pattern held once we started having this conversation with prospective customers directly. Explaining the philosophy behind cost-per-resolution pricing, on its own, produced polite nodding. Showing a prospect the equivalent chart, built from a short pilot period on their own ticket data, produced the kind of questions that actually move a deal forward — mostly about which of their own ticket types were driving the disproportionate cost, which is a far more productive conversation than any pricing-model discussion in the abstract.

If you're trying to make a similar change inside your own organization, or trying to sell a similar pricing model to a skeptical buyer, our honest recommendation is to skip the philosophical argument first and go straight to the chart. The number that actually moves people is rarely the argument for why a metric is theoretically wrong — it's the concrete evidence that the theoretically-wrong metric is already costing something specific, today, and getting worse.

We've since made this chart part of every quarterly business review internally, specifically so the argument never has to be re-made from scratch — it's now just a standing number people expect to see, the same way they'd expect to see revenue or churn, which has done more to keep the pricing model honest over time than any single policy decision could.

In practice

If you're pricing something similar, the concrete test we'd suggest: take your current pricing unit and ask what behavior it rewards if a customer optimizes purely for that unit. If the answer is "more messages" or "more calls" rather than "more problems that stop existing," the unit is measuring the wrong thing, regardless of how standard it is in the category.