Why this nixos-tools pattern gets slow in production: a source-level explanation
Some background first. Our setup is nixos-tools plus three downstream services, seven figures of daily requests, peaking around nine in the evening.
We also fixed monitoring along the way: replaced average-based alerts with percentiles and split them per endpoint. False alerts dropped by about seventy percent and the on-call rotation visibly cheered up.
POST /api/uploads → CDN origin pull