Capacitor · Ops · Open-sourced a rate-limiting middleware for capacitor-ops with token bucket and sliding window

696
CAr/capacitor-ops·posted by chen_dev·yesterdayTranslation

Open-sourced a rate-limiting middleware for capacitor-ops with token bucket and sliding window

Short version: capacitor-ops needs almost no tuning at small and medium scale — the point where it starts to hurt is much further out than most people assume. Full measurements below.

// Minimal reproduction: you must use a real long-tail distribution here.
// Uniform load-test traffic will never trigger this.
func (s *Server) handle(ctx context.Context) error {
    conn, err := s.pool.Acquire(ctx)
    if err != nil {
        return fmt.Errorf("acquire: %w", err)
    }
    defer conn.Release()

    return s.do(ctx, conn)
}

Worth noting: the official docs do cover this, just in a very inconspicuous spot. I only found it reading the source comments, where the author explains the reasoning — roughly "so that it degrades into predictable behaviour in extreme cases".

One last trap: in container environments remember to adjust the memory-related parameters in step. Otherwise the host limit and the process expectation disagree, and the symptom is intermittent, unreproducible failure.

We also fixed monitoring along the way: replaced average-based alerts with percentiles and split them per endpoint. False alerts dropped by about seventy percent and the on-call rotation visibly cheered up.

What genuinely surprised me was the tail. The average looked great while P99 jumped by an order of magnitude past some threshold. The cause was not capacitor-ops itself but our upstream connection reuse — the load test traffic was too clean and hid the long-tail requests.

194 comments

194 comments

· first 120 loaded
M
Cchen_dev·2 days ago

I just read the capacitor-ops source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

512
Rrase·5 hours ago

I just read the capacitor-ops source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

508
Cchen_dev·28 minutes ago

One counter-example: below capacitor-ops 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

468
Bbob_chen·1 hour ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

428
Hhuang_ke·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

466
Kkite·2 days ago

One counter-example: below capacitor-ops 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

309
Rrase·yesterday

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

1
Zzhou_yi·just now

Saved. I am reworking this area this week — this saves a lot of wrong turns.

64
Nnikic·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

496
Rrase·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

56
Kkernel_panicMod·5 hours ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

468
Aalice_dev·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

261
Lli_ming·2 days agoLevel 6

Saved. I am reworking this area this week — this saves a lot of wrong turns.

5
Lli_ming·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

254
LlinlinOP·2 days ago

This is not a capacitor-ops problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

136
Zzhou_yi·2 days agoLevel 6

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

103
Ddev_zhou·2 days agoLevel 6

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

101
Llinlin·2 days agoLevel 6

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

4
Ttang_hao·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

54
Aalice_dev·2 days agoLevel 6

I just read the capacitor-ops source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

326
Lli_mingOP·just now

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

26
Zzhu_zong·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

153
Zzhu_zong·2 days agoeditedLevel 6

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

429
Kkernel_panic·5 hours agoLevel 6

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

416
Ttang_haoOP·2 days agoedited

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

3
Lli_ming·1 hour ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

40
Sslow_query·2 days agoLevel 6

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

276
Cchen_dev·12 minutes agoLevel 6

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

83
Sslow_query·1 hour agoLevel 6

This is not a capacitor-ops problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

13
Rran_bo·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

3
Llinlin·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

57
Rrase·2 days agoedited

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

40
Cchen_devOP·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

202
Rran_bo·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

367
Oops_wang·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

338
Rran_bo·2 days agoeditedLevel 6

This is not a capacitor-ops problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

129
Aalice_dev·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

1
Llinlin·5 hours ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

1
Sswoole_leeOP·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

405
Ttang_hao·2 days agoedited

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

427
Ttang_hao·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

421
Zzhu_zongOP·28 minutes ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

354
Rran_bo·1 hour ago

I just read the capacitor-ops source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

71
Oops_wangOP·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

506
Kkernel_panic·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

156
Hhuang_ke·2 days agoedited

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

466
Kkernel_panic·5 hours ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

177
Kkite·12 minutes ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

92
Ddev_zhou·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

1
Kkernel_panicOP·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

68
Zzhou_yi·1 hour ago

One counter-example: below capacitor-ops 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

63
Sswoole_lee·3 minutes ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

10
Ttang_hao·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

7
Kkite·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

455
Ttang_hao·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

318
Mmike_xu·1 hour ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

1
Wwinter·2 days agoedited

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

360
Kkernel_panic·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

347
Zzhu_zong·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

305
Ddev_zhou·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

285
Zzhu_zongMod·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

247
Ddev_zhou·5 hours ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

233
Rran_bo·5 hours ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

30
Ttang_hao·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

220
Oops_wang·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

7
Oops_wang·1 hour ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

216
Ddev_zhou·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

205
Zzhou_yi·2 hours agoedited

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

164
Cchen_dev·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

286
Zzhu_zong·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

70
Sslow_query·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

300
Cchen_dev·2 hours ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

8
Mmike_xu·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

39
Ddev_zhouOP·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

450
Aalice_devOP·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

245
Bbob_chen·2 days ago

One counter-example: below capacitor-ops 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

64
Ddev_zhou·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

4
Llinlin·2 days agoedited

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

20
Bbob_chen·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

3
Oops_wang·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

2
Sswoole_lee·2 days ago

One counter-example: below capacitor-ops 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

123
Hhuang_keOP·3 minutes ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

366
Cchen_dev·2 days ago

I just read the capacitor-ops source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

25
Ddev_zhou·2 days agoLevel 6

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

4
Rran_bo·2 days agoedited

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

102
Kkernel_panic·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

160
Ttang_haoOP·2 days agoedited

I just read the capacitor-ops source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

77
Mmike_xu·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

7
Rran_bo·3 minutes ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

47
Nnikic·2 days ago

This is not a capacitor-ops problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

128
Kkernel_panic·3 minutes agoedited

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

168
Hhuang_ke·2 days ago

One counter-example: below capacitor-ops 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

1
Lli_ming·2 days ago

This is not a capacitor-ops problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

10
Lli_ming·2 days ago

I just read the capacitor-ops source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

171
Aalice_dev·2 days ago

One counter-example: below capacitor-ops 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

127
Zzhou_yi·2 days ago

One counter-example: below capacitor-ops 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

163
Lli_ming·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

3
Kkite·12 minutes ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

121
Sslow_query·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

279
Kkite·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

379
Mmike_xu·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

86
Bbob_chenOP·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

50
Ddev_zhouOP·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

46
Aalice_dev·2 days ago

This is not a capacitor-ops problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

87
Aalice_dev·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

78
Ttang_hao·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

83
Lli_ming·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

77
Kkernel_panic·2 days agoedited

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

72
Ttang_hao·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

72
Nnikic·yesterday

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

65
Zzhou_yi·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

65
Rran_bo·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

55
Wwinter·just now

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

41
Rrase·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

18
Rrase·1 hour ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

13
Mmike_xu·12 minutes ago

This is not a capacitor-ops problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

7
Llinlin·2 days ago

This is not a capacitor-ops problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

173
Sslow_query·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

1
Rran_bo·2 days agoedited

I just read the capacitor-ops source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

7
Sswoole_lee·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

136

This is the post detail page /en/c/capacitor-ops/post/p12. Posts and comments are generated deterministically from a seeded PRNG, so the same post always renders the same content and the link can be shared, reloaded and indexed. In production this page reads MySQL for the post, Redis for hot-post caching, and fetches the whole comment tree in a single query on the path column.

See the database schema →