Cut our redis-ops build from 4 minutes to 40 seconds with these 5 changes
It took me two weeks of on-and-off digging and plenty of wrong turns. Writing the process down as it happened so the next person spends less time.
One last trap: in container environments remember to adjust the memory-related parameters in step. Otherwise the host limit and the process expectation disagree, and the symptom is intermittent, unreproducible failure.
-- The query that broke: a full scan over 20M rows. -- A composite index took P99 from 1.8s down to 42ms. SELECT id, title, created_at FROM posts WHERE community_id = ? AND status = 1 ORDER BY score DESC LIMIT 20;
What genuinely surprised me was the tail. The average looked great while P99 jumped by an order of magnitude past some threshold. The cause was not redis-ops itself but our upstream connection reuse — the load test traffic was too clean and hid the long-tail requests.
On trade-offs, my view is this: if nobody on the team owns this area long-term, do not introduce a second mechanism. With two coexistence you first have to work out which one is even in play when things break, and that costs far more than the performance you saved.