Overview
by Julia Berg
No spike in CPU. Error budgets looked fine at a glance.
Users still reported blank pages in a thin slice of traffic.
Logs only showed upstream resets with no clear application exception.
The culprit was a stale keep-alive timeout between nginx and the app.
Aligning idle timeouts stopped the intermittent 502s within an hour.
We also added a dashboard for upstream reset reasons so the next page is faster.