All articlesEngineering

How we load-test an AI Agent before a Black Friday spike

Marcus Ibe10 min read
Engineeringlimit

Our busiest customers see traffic spike 6 to 8x on Black Friday, and every one of those spikes arrives as a wall of concurrent conversations, not a gradual ramp. An Agent that handles 50 conversations per minute comfortably can fall over at 400 if nothing was tested ahead of time, and the load-testing process we run every year exists specifically to find that failure point before a real customer does.

Why Black Friday is different

3s limitnormalBlack Friday

The thing that makes Black Friday traffic distinct from an ordinary busy day isn't just the volume — it's the shape of the arrival curve. A gradual increase over hours gives every part of a system, human and automated, time to adapt. A spike that arrives within minutes of a sale going live gives none of that runway, and a system that's never been tested against that specific shape of load can have failure modes that never show up under gradual growth, because gradual growth never actually stresses concurrency the same way.

That distinction is why we don't consider ordinary peak-hour monitoring sufficient preparation. Normal peak hours, even busy ones, still arrive gradually enough that a system under strain has time to shed load gracefully or for a human to notice and intervene. A genuine spike doesn't give you that time, which is exactly why it needs to be tested deliberately rather than assumed to be covered by ordinary monitoring.

How we simulate the spike

We simulate the spike two weeks out: replay a sample of real conversation patterns from the previous year's peak day, at increasing concurrency, until response latency crosses our 3-second threshold. That number becomes the capacity ceiling we plan around. Using real historical conversation patterns rather than synthetic test traffic matters specifically because real conversations have a distribution of complexity — some resolve in two messages, some need several tool calls — and a synthetic load test built from uniform, simple test cases would miss exactly the kind of complexity-driven bottleneck that a real spike actually exposes.

The two-week timing isn't arbitrary either. It's early enough that whatever bottleneck the test surfaces can actually be fixed before the real event, but late enough that the infrastructure being tested reflects what will actually be running on the day itself, rather than an earlier configuration that might have since changed.

What the bottleneck actually is

The bottleneck is almost never the model itself — it's usually a synchronous call to a slow downstream system, like an order lookup API that wasn't built for this kind of concurrency. Load testing surfaces those dependencies before customers do, and the fix, once found, is almost always more mundane than anyone expects going in: adding a connection pool, making a call asynchronous where it didn't need to block, or adding a cache in front of a downstream system that gets hit with the exact same query repeatedly during a spike in a way it never does under normal traffic.

One specific pattern we've seen repeatedly: a downstream inventory or order API that performs perfectly well under normal load starts queuing requests once concurrent calls cross some threshold specific to that system's own infrastructure — a threshold that has nothing to do with the Agent's own capacity and everything to do with a dependency the team building the Agent doesn't directly control. Finding that threshold two weeks before Black Friday, in a controlled test, is a very different experience than finding it live, on the day, with real customers waiting.

The teams that skip this find out the hard way at 9am on Black Friday. The ones that don't spend an afternoon two weeks earlier instead — and the asymmetry between those two outcomes, a planned afternoon of testing versus an unplanned, high-stakes emergency, is the entire case for doing this every single year rather than assuming last year's capacity ceiling still holds.

Common questions

Does the capacity ceiling found in testing actually hold up on the real day, or does real traffic still surprise you? It holds up closely in most years, though we always keep some margin above the tested ceiling rather than planning right up against it, specifically because real traffic occasionally has a wrinkle a simulation based on last year's patterns didn't anticipate — a new promotion structure, a product launch timed to coincide with the sale, that kind of variable.

How do you decide the 3-second latency threshold specifically, rather than some other number? It's based on where customer abandonment during a conversation starts climbing meaningfully in our own data — below that threshold, most customers wait without noticing; above it, a rising share start abandoning the conversation before getting an answer, which is the actual behavior we're trying to prevent rather than an arbitrary technical target.

The year we found the bottleneck too late

It's worth describing the year this process didn't work as well as it should have, because the failure taught us more about how to run the test properly than any of the years it went smoothly.

Two years ago, our load test two weeks before Black Friday found a capacity ceiling comfortably above what we expected to need, based on the previous year's peak traffic plus a reasonable growth margin. The test passed. We didn't dig further, because there was no obvious reason to.

On the day itself, one specific customer's traffic came in well above their own historical pattern, driven by an unusually successful promotional email campaign nobody on our side had known was planned, layered on top of the general Black Friday spike everyone was already expecting. The combined load crossed our tested ceiling by a meaningful margin, and the specific downstream dependency that buckled was the same category of bottleneck we've described generally — a synchronous order-lookup call that started queuing once concurrent requests crossed a threshold the load test's traffic pattern, built from the previous year's data, had never actually reached.

The fix that day was manual and reactive: a temporary, more aggressive routing-to-human threshold, deployed live, that shed load by handing off more conversations than usual rather than letting response times climb past the point where customers started abandoning. It worked, but it worked because someone was watching closely enough to catch the pattern in real time and had the authority to make a live configuration change under pressure — not because the system was built to handle it gracefully on its own.

The change we made afterward wasn't just "increase the tested ceiling with more margin," though we did that too. It was building a live headroom alert — a standing check, active only during the specific high-traffic window around Black Friday, that flags when actual concurrent load crosses a percentage of the last tested ceiling, well before hitting the ceiling itself, giving the team time to react deliberately rather than discovering the problem only once customers were already experiencing it.

We also added a specific step to the load-testing process itself: rather than testing only against the previous year's exact traffic pattern, we now test against that pattern scaled up by a margin meant to represent an unplanned surprise on top of expected growth — not because we can predict what the surprise will be, but because the specific year we got caught taught us that assuming no surprise is itself the riskiest assumption in the whole process.

That change came directly out of a bad experience, and we mention it specifically because it's the kind of lesson that's much cheaper to learn from someone else's account of it than from your own live incident — which is exactly why we're including it here rather than only describing the years everything went according to plan.

Common questions

How far in advance do you need to know about an unusual promotional campaign, like the one that caused the near-miss described in this piece, for it to actually get folded into load testing? Ideally at least the same two weeks the load test itself runs before the event, though even a few days' notice is enough to add a rough traffic-scaling adjustment to the test rather than running against an unadjusted baseline.

Does load testing ever reveal a problem that isn't actually about capacity at all, but about correctness under concurrent load — for instance, a race condition that only shows up under heavy simultaneous traffic? Yes, and this is one of the more valuable, if less expected, things load testing catches — concurrency bugs that never surface under normal traffic levels but appear reliably once enough simultaneous requests are hitting the same code path, which is a correctness issue load testing surfaces as a side effect of testing for capacity.

How do you decide the specific traffic multiplier to test against, beyond "last year's peak plus some margin"? We look at year-over-year growth trends for the specific customer base being tested, plus the surprise-margin adjustment described above, rather than using a single fixed multiplier across every customer regardless of their own specific growth trajectory.

Is the live headroom alert described above specific to Black Friday, or does it run year-round as a general safeguard? It's active specifically during defined high-traffic windows, Black Friday being the largest but not the only one, rather than running as a constant alert year-round, since a threshold tuned for an expected spike period would generate unnecessary noise during ordinary traffic the rest of the year.

What's the actual cost of running this load test every year, in terms of engineering time, and is it worth it relative to the risk it addresses? It takes a focused afternoon of setup and monitoring for most deployments, which is a small cost relative to the risk of an uncontrolled capacity failure during the single highest-stakes traffic period of the year for most of the customers this applies to — the asymmetry between the cost of testing and the cost of not testing is, in our view, not close.

What happens if the load test finds a bottleneck that can't realistically be fixed in the two weeks before the event? We plan a fallback for that specific scenario in advance — usually a more aggressive routing-to-human policy that kicks in automatically once concurrency crosses the known bottleneck point, so the system degrades gracefully into a slower, more human-assisted mode rather than failing outright.

Do smaller stores without Black-Friday-scale traffic need this kind of testing too? The same principle applies at smaller scale for any predictable spike — a smaller store's version of this might be a single major promotional email rather than a national shopping holiday, but the same test-before-the-spike discipline applies regardless of the absolute traffic numbers involved.

The near-miss described above is the example we now use internally whenever a newer team member asks why we treat a passed load test with a comfortable margin as provisional rather than final. A test that passed against last year's traffic pattern only tells you the system can handle last year's traffic pattern plus your chosen margin — it says nothing about a genuinely new variable, like an unplanned promotional campaign, that no historical pattern could have anticipated, which is exactly why the surprise-margin adjustment and the live headroom alert both exist as permanent parts of the process now, not optional extras.

A specific detail from the near-miss year worth adding: the promotional campaign that caused the unexpected surge wasn't something the customer had deliberately hidden from us — it simply hadn't occurred to anyone on either side that a marketing calendar detail was relevant information to share with the team running load tests. That's led to a small but genuinely useful process change: we now explicitly ask every customer, as part of scheduling their annual load test, whether anything unusual is planned around the same period, rather than assuming we'd naturally be told about anything relevant.