No spike in CPU. Error budgets looked fine at a glance. Users still reported blank pages in a thin slice of traffic. Logs only showed upstream resets with no clear application exception. The culprit was a stale keep-alive timeout between nginx and the app. Aligning idle timeouts stopped the intermittent 502s within an hour. We also added a dashboard for upstream reset reasons so the next page is faster.
We extracted one high-churn billing endpoint behind a strangler facade. Dual-writes ran for two weeks while we compared totals nightly. A feature flag controlled read traffic so we could roll back instantly. The hardest part was matching edge-case rounding in legacy invoices. Cutover finished with no customer-facing downtime and a smaller blast radius. We kept the facade until three more endpoints followed the same path.
We needed locks so queue workers did not process the same job twice. Redis SET NX was faster under load and easy to expire automatically. Postgres advisory locks were simpler operationally for our small team. Failover behavior mattered more than raw latency in our case. We chose Postgres first, then moved hot paths to Redis later. Pick the lock store you can operate confidently at 3am.