Open-sourced a rate-limiting middleware for high-concurrency-review with token bucket and sliding window
It took me two weeks of on-and-off digging and plenty of wrong turns. Writing the process down as it happened so the next person spends less time.
One last trap: in container environments remember to adjust the memory-related parameters in step. Otherwise the host limit and the process expectation disagree, and the symptom is intermittent, unreproducible failure.
-- The query that broke: a full scan over 20M rows. -- A composite index took P99 from 1.8s down to 42ms. SELECT id, title, created_at FROM posts WHERE community_id = ? AND status = 1 ORDER BY score DESC LIMIT 20;
Worth noting: the official docs do cover this, just in a very inconspicuous spot. I only found it reading the source comments, where the author explains the reasoning — roughly "so that it degrades into predictable behaviour in extreme cases".
We also fixed monitoring along the way: replaced average-based alerts with percentiles and split them per endpoint. False alerts dropped by about seventy percent and the on-call rotation visibly cheered up.