Policy engines, not smarter prompts: what actually keeps an agent safe
Every time we upgraded to a smarter model, we expected fewer mistakes. What we actually got, more often than not, was more confident mistakes — the same wrong answer, delivered with less hedging and more fluency. That single observation reshaped how we think about safety in a support agent, enough that we ended up walking the whole question through a structure borrowed from advertising research rather than machine learning: awareness that something is off, comprehension of the actual mechanism, conviction once we had numbers instead of a hunch, and the concrete actions that followed from all three.
None of what follows is a claim that models don't matter. It's a claim that model quality solves a narrower slice of the safety problem than most teams assume, and that the slice it doesn't solve — deciding whether the agent should have the authority to act at all — needs its own layer, built independently of which model happens to be running that day.
Awareness
The first sign wasn't a dramatic failure, the kind that shows up in a postmortem with a clear before-and-after. It was a pattern buried in the weekly transcript review: a model that hallucinates a refund status with total conviction reads, to the customer receiving it, exactly the same as a correct answer — right up until it turns out to be wrong, usually days later when the actual refund never arrives. A weaker, cheaper model that says "I'm not sure, let me check" is strictly safer than a stronger one that guesses fluently, and once we started grading transcripts for that distinction specifically instead of folding it into a general quality score, the gap became too obvious to keep treating as noise.
It also showed up in something more mundane: how often an upgrade to a newer model changed the actual outcome of borderline cases versus how often it just changed the wording. We compared a batch of ambiguous refund requests across two model versions and found the newer model was, if anything, slightly more likely to find a plausible-sounding justification for approving something the older model had flagged for review. Fluency and judgment aren't the same capability, and a model that got better at the first didn't automatically get better at the second.
That's what made this an awareness problem rather than a one-off bug. Nobody had decided to trade judgment for fluency. It happened by default, because every benchmark we were using to pick models measured how good the answer sounded, not how well the system knew when not to answer at all.
Comprehension
Once we went looking for the actual mechanism, it wasn't subtle — it was just missing a step nobody had named. Good agent planning breaks into six distinct stages: classify the intent, check whether the agent even has the authority to act on it, fetch the minimum context needed, propose an action, validate that action against policy, then execute. Written out like that, it looks almost too obvious to be the fix. What we found, going back through our own earlier implementation and a few others we looked at for comparison, is that almost every "agent" collapses stages two and five into the model's own judgment instead of treating them as separate, auditable checks.
Stages one and three are the ones every agent gets right early, because they're the ones a language model is naturally good at: understanding what the customer wants, and pulling the context needed to address it. Stage four — proposing an action — is also something models do well, often too well, because proposing a plausible-sounding action is exactly the skill a fluent model has been optimized for. The failure lives specifically in the gap between four and six: nothing was checking whether the agent had the standing to make that proposal real, and nothing was checking the proposal against a rule that didn't care how confident the phrasing was.
That's the part smarter models don't fix, and can't fix by construction. A better model gets measurably better at stage four — the proposals get more plausible, the language gets more natural, the edge cases get handled with more nuance. It doesn't get better at stage two, whether it should have the authority to act at all, because that's not a language modeling problem in the first place. It's a systems design gap sitting one layer above the model, and no amount of additional training data teaches a system a boundary that was never encoded anywhere in its architecture.
Put concretely: "the model decided to refund $400" is a sentence that should worry anyone reviewing it, no matter which model produced it. "The model proposed a refund, a policy engine checked eligibility and the dollar limit, then it executed" describes the identical customer-facing outcome with an entirely different risk profile, because the authority to act moved to a place that isn't trying to sound helpful while it makes the call.
Conviction
We didn't take any of this on faith once we understood the mechanism — we built the missing layer and measured its actual effect, on the same evaluation set, before and after. Overall resolution rate barely moved, which on its own would read as a null result if that's the only number you tracked. It's exactly the number most teams would report and stop there.
The number that moved was escalations for the wrong reason — cases the agent should have handled confidently but bounced to a human out of miscalibrated caution, or, more worryingly, cases it shouldn't have touched at all but did anyway before the policy layer existed. That rate dropped by more than half. Resolution rate answers "did a human get pulled in." Wrong-reason escalation rate answers something closer to "was the system's judgment about its own authority actually correct," and that second number is the one that predicts whether you can trust the agent with a harder case next month, not just this one.
Two smaller guardrails earned their place the same way — not because they sounded prudent, but because we could reproduce the specific failure each one prevented. Every retryable action needs an idempotency key, not just a "try again" loop layered on top of it. Without one, a network timeout on a refund call that had actually already succeeded turns a naive retry into a second refund, and we found exactly this failure mode in a load test before it ever reached a real customer — once we went looking for it deliberately instead of assuming timeouts were harmless.
The other was a rate cap, tested the same way: an agent that can issue unlimited refunds per hour is one bad reasoning loop away from an expensive afternoon, not one malicious actor away from it. We simulated a scenario where a subtle bug caused the agent to repeatedly re-approve a borderline case across a burst of retries, and a simple per-action, per-time-window cap stopped the bleeding within seconds instead of requiring a human to notice and intervene manually hours later.
Action
What we actually shipped, in order of how load-bearing each piece turned out to be: a policy engine sitting between "the model proposed X" and "X happened," checking eligibility, dollar limits, and standing authority on every write action, independent of which model generated the proposal. Idempotency keys on every retryable call, no exceptions, including the ones that felt too simple to need one. Per-action, per-window rate caps, sized conservatively at first and loosened only after watching real traffic against them for a few weeks.
Alongside that, a scope review on every tool the agent gets access to, done before the tool ships rather than after an incident: a support agent doesn't need write access to pricing, inventory, or admin settings, even if the model "probably wouldn't" misuse it, because "probably" was never treated as a security boundary anywhere else in the stack and there's no reason to make an exception for this one.
The practical version of this, if you're building something similar: before any new write-capable tool goes live, write down in one sentence what the worst plausible misuse of it looks like, then check whether the policy engine — not the prompt, not the model's own judgment — actually prevents that specific misuse. If the honest answer is "the model probably wouldn't do that," the tool isn't ready, regardless of how well it performed in testing.
Where this breaks down
None of this is free, and it's worth being honest about the cost side rather than only the benefit side. A policy engine adds a real check on the critical path of every action, which means added latency, however small, on every write operation — noticeable enough that we had to optimize the eligibility lookups separately once volume grew past a few thousand actions a day. It also means a second place things can be wrong: a policy rule that's stricter than it needs to be silently creates unnecessary escalations, and unlike a model mistake, a bad policy rule doesn't show up as an obviously wrong answer — it shows up as a quiet increase in cases routed to humans for no real reason, easy to miss unless you're specifically watching for it.
Common questions
Doesn't a policy engine just move the risk from the model to whoever writes the policy rules? Yes, and that's the point — it moves the risk somewhere auditable. A policy rule is a piece of code or config that a human wrote, reviewed, and can point to when asked why a specific action was allowed. A model's judgment in the moment is none of those things, even when the model happens to get the answer right. Moving risk from an opaque place to a visible one is the actual improvement, not a sleight of hand.
How much of this can you get from a good system prompt instead of a separate engine? Less than it looks like you can. A system prompt is a request, and the model can be talked out of a request through a long enough conversation, an unusual phrasing, or content injected from outside the conversation entirely. A policy engine is a check that runs regardless of what the model was told or how it was persuaded, which is the actual property you need for anything touching money or account access.
What's the smallest version of this worth building, if you're just starting out? A single eligibility check and a hard dollar cap on the riskiest action type, enforced outside the model, before anything else. That alone closes the two failure modes most likely to actually happen early — an eligible-looking but ineligible refund, and an amount nobody would have approved by hand. Everything else here is worth building once that exists and you've watched it in production for a few weeks.
Does adding a policy engine slow down how fast you can ship new agent capabilities? It slows down shipping anything that touches money or account state, on purpose, and that's a trade we'd make again. It doesn't slow down the parts of the agent that only read and explain — those never needed the same check in the first place, and treating them the same way would be its own kind of over-engineering.
The one-page version
If you're building this from scratch and want the condensed version rather than the full argument: every write action gets a policy check that runs independently of the model, before execution, not as a suggestion the model can override. Every retryable call gets an idempotency key. Every action type gets a rate cap sized to what a human would approve without hesitation, reviewed and adjusted after watching real traffic. Every tool the agent can call gets a one-sentence answer to "what's the worst misuse of this with bad input" written down before it ships.
The order we'd build these in, if starting today: the policy engine first, because it's the piece that actually changes what "the agent decided X" means from a liability into an audited action. Idempotency second, because it's cheap to add early and expensive to retrofit once retries are already happening in production without it. Rate caps third, sized conservatively and loosened deliberately rather than left uncapped by default. The tool-scoping habit last, not because it matters least, but because it's a practice that has to become routine rather than a one-time audit — it only works if it's applied to the next tool, and the next one after that, indefinitely.
The single test we'd ask any team to run on their own agent before trusting it with real money: pick the riskiest action it can take, and ask whether the sentence describing how that action gets approved contains the word "model" anywhere before the word "policy" or "limit." If it does, the authority hasn't actually moved yet, no matter how good the surrounding prompt sounds.
The other honest limitation is ownership. A model is something you can retrain, swap, or roll back with a version number. A policy engine is a living set of business rules that someone has to actually maintain as the business changes — a new discount program, a new market with different return windows, a new payment provider with different refund semantics all require someone to update the policy layer, and if that update lags the product change, the policy engine becomes the thing blocking legitimate actions instead of the thing preventing bad ones. Treating it as a one-time build instead of an owned, versioned piece of the system is the most common way we've seen teams undo the benefit of building it in the first place.
None of this makes the agent seem smarter in a demo, and that's roughly the point. It makes the agent boring in exactly the way that matters for something handling real money and real customers: predictable under pressure, auditable after the fact, and safe to hand real authority to — because the authority was never actually the model's to give away in the first place.