Two weeks with verilog-newbie: what feels right and what drives me up the wall
Some background first. Our setup is verilog-newbie plus three downstream services, seven figures of daily requests, peaking around nine in the evening.
We also fixed monitoring along the way: replaced average-based alerts with percentiles and split them per endpoint. False alerts dropped by about seventy percent and the on-call rotation visibly cheered up.
Worth noting: the official docs do cover this, just in a very inconspicuous spot. I only found it reading the source comments, where the author explains the reasoning — roughly "so that it degrades into predictable behaviour in extreme cases".