Why this monitoring-pro pattern gets slow in production: a source-level explanation
Most monitoring-pro articles stop at "how to use it" and never cover "when not to use it". This is an attempt at the second half.
Worth noting: the official docs do cover this, just in a very inconspicuous spot. I only found it reading the source comments, where the author explains the reasoning — roughly "so that it degrades into predictable behaviour in extreme cases".
-- The query that broke: a full scan over 20M rows. -- A composite index took P99 from 1.8s down to 42ms. SELECT id, title, created_at FROM posts WHERE community_id = ? AND status = 1 ORDER BY score DESC LIMIT 20;
What genuinely surprised me was the tail. The average looked great while P99 jumped by an order of magnitude past some threshold. The cause was not monitoring-pro itself but our upstream connection reuse — the load test traffic was too clean and hid the long-tail requests.
One last trap: in container environments remember to adjust the memory-related parameters in step. Otherwise the host limit and the process expectation disagree, and the symptom is intermittent, unreproducible failure.