Postmortem: how we exhausted the connection pool in edge-computing-show
Some background first. Our setup is edge-computing-show plus three downstream services, seven figures of daily requests, peaking around nine in the evening.
We also fixed monitoring along the way: replaced average-based alerts with percentiles and split them per endpoint. False alerts dropped by about seventy percent and the on-call rotation visibly cheered up.
1637 votes total