Published with customer names, hostnames and internal service names removed. Everything else is as filed on 15 May 2026, including the parts that make us look slow.
Times are UTC. Owners are initials. I am DW.
There are initials in the timeline because someone has to be findable afterwards, and no initials in the root cause because a person is not a root cause. If the answer to "why did this happen" is a name, the investigation stopped early. This is the eleventh of these I have written. The table is the only part that gets easier.
Forty three minutes, 09:19:40 to 10:02:50 on Thursday 14 May 2026. The public card gateway returned 502 for 24 of those minutes at 100%, and at between 6% and 98% for the rest. At a steady 780 requests per second, about 1.45 million requests failed.
No money moved incorrectly. The gateway fails ahead of the authorisation call and the idempotency layer sits behind it, so nothing was double charged and nothing was captured twice. Merchants retrying during the window got 502 as well, which is the correct outcome and does not feel like one.
Two merchants missed their settlement batch window and were replayed by hand on 15 May.
| Time | Event | Who saw it |
|---|---|---|
| 09:11:02 | Canary of gateway 2026.5.14-3 to 1 of 34 pods. Canary gate green after 4 minutes. | pipeline |
| 09:14:30 | Rollout to the remaining 33 pods, 6 at a time. Completes 09:18:10. | pipeline |
| 09:19:20 | A merchant's daily retry batch starts. Its reference field ends in base64 padding. | nobody |
| 09:19:40 | First 502. Two pods stop answering readiness probes. | nobody |
| 09:21:05 | Autoscaler goes 34 to 41 pods on CPU. | nobody |
| 09:22:40 | 502 rate at 6%. Liveness starts restarting pods. | nobody |
| 09:23:31 | Support engineer posts "merchant dashboard is throwing errors" in the help channel. | RK |
| 09:25:10 | Autoscaler at 60, the configured maximum. New pods take traffic and stop answering within about 90 seconds. | nobody |
| 09:26:02 | DW starts looking. Latency panel shows p99 at 2.500s, flat. Read as slow. | DW |
| 09:28:44 | 502 rate at 100%. | DW |
| 09:30:12 | DW ties it to the 09:14 rollout. Decision to roll back. | DW |
| 09:31:50 | GatewayErrorRatioHigh pages. The condition had been continuously true since 09:20:50. | pager |
| 09:33:05 | Deploy CLI returns 502. It fetches its token through the gateway. | DW |
| 09:36:20 | Break-glass path located in the runbook. It needs a credential nobody on the call held. | DW, MO |
| 09:41:10 | MO has the credential. kubectl rollout undo issued. | MO |
| 09:43:20 | One pod held on the new image, pulled out of the load balancer, for later analysis. | DW |
| 09:44:00 | First pods back on 2026.5.14-2. They stay up. | MO |
| 09:52:15 | 24 of 60 pods on the old image. 502 rate 41%. | MO |
| 09:58:30 | All pods on the old image. | MO |
| 10:02:50 | 502 rate at baseline. Mitigated. | DW |
| 10:19:00 | Autoscaler back to 35 pods. | nobody |
| 11:48:00 | CPU profile from the held pod: 96% of samples inside a single regular expression match. | AF |
A validation pattern added to the shared request middleware nests two unbounded quantifiers over overlapping character classes, so rejecting a value costs roughly twice as much for every additional allowed character that precedes the disallowed one, and one merchant's reference format supplies between 34 and 46 of them.
^([A-Za-z0-9]+[-_ ]?)+$
The group takes one or more alphanumerics followed by an optional separator. The outer + repeats the group. For a value that is entirely alphanumeric, the number of ways to divide it between repetitions of that group is the number of ordered compositions of its length, . When the value matches, the engine finds one of those divisions immediately and stops. When it does not match, because of a single = at the end, the engine has to try all of them before it can say no. At 32 characters that is divisions, a little over two billion, which is why the number below reads in seconds.
Measured on my laptop, one core, node, on a string of n alphanumerics followed by one =:
| n | time to fail |
|---|---|
| 24 | 99 ms |
| 26 | 398 ms |
| 28 | 1.62 s |
| 30 | 6.46 s |
| 32 | 26.1 s |
| 40 | 1.9 hours, extrapolated |
| 46 | 5 days, extrapolated |
V8 offers no way to abandon a match in progress. Once a worker enters that call, the worker is gone until the process is killed.
The pattern was reviewed by two engineers and shipped with eleven unit tests. All eleven passed in under a millisecond, because all eleven inputs matched. Nobody writes a unit test for a value that gets rejected slowly. That is a property of tests, not of the people who wrote them.
/healthz. No path through the process skipped it.The latency panel. The gateway histogram's largest finite bucket boundary is 2.5 seconds, and when a quantile falls in the +Inf bucket, histogram_quantile returns the upper bound of the second highest bucket. That is documented behaviour and it is exactly what happened: the panel read 2.500 for forty three minutes. I looked at it at 09:26 and concluded we were slow. We were not slow, we were stopped. The rate of the +Inf bucket went from about 4 per minute to 61,000 and there was no panel for it.
The alert. GatewayErrorRatioHigh carries for: 10m, and Alertmanager's default group_wait is 30 seconds. The condition became true at 09:20:50 and the page landed at 09:31:50. A support engineer and a dashboard both beat it. The for: 10m was set in 2023 after a month of flapping, which was reasonable then and was still there in May 2026, which is the actual problem.
The logs. We log on response. A request that never responds writes nothing, so the last line from a stalled worker is the line before the middleware. Log volume from the gateway went down for forty three minutes, which is the opposite of what four people were grepping for.
The runbook. Step 2 is "check recent deploys" and links to a board that lists image deploys sorted ascending by start time. The 09:14 rollout was at the bottom of a scrolling list.
Dates are commitments. Two action items from a postmortem in 2024 quietly expired in a tracker, so they live in the document now.
+Inf rate panel placed next to every p99 panel we own. RK, 21 May.GatewayErrorRatioHigh moved to for: 2m on a one minute window, plus a new page at 5xx above 25% with for: 0m. DW, 21 May.recheck. MO, 12 June.