Modding · Reading list · Interview notes: how I answered a question about the modding-books concurrency model

2.2K
MOr/modding-books·posted by ops_wang·3 days agoTooling

Interview notes: how I answered a question about the modding-books concurrency model

Most modding-books articles stop at "how to use it" and never cover "when not to use it". This is an attempt at the second half.

-- The query that broke: a full scan over 20M rows.
-- A composite index took P99 from 1.8s down to 42ms.
SELECT id, title, created_at
  FROM posts
 WHERE community_id = ?
   AND status = 1
 ORDER BY score DESC
 LIMIT 20;

We also fixed monitoring along the way: replaced average-based alerts with percentiles and split them per endpoint. False alerts dropped by about seventy percent and the on-call rotation visibly cheered up.

On trade-offs, my view is this: if nobody on the team owns this area long-term, do not introduce a second mechanism. With two coexistence you first have to work out which one is even in play when things break, and that costs far more than the performance you saved.

What genuinely surprised me was the tail. The average looked great while P99 jumped by an order of magnitude past some threshold. The cause was not modding-books itself but our upstream connection reuse — the load test traffic was too clean and hid the long-tail requests.

1015 comments

1015 comments

· first 120 loaded
M
Zzhu_zongOP·2 hours ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

471
Wwinter·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

499
Sslow_query·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

462
Hhuang_ke·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

476
Sswoole_lee·2 days agoedited

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

115
Bbob_chen·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

76
Aalice_dev·2 days ago

One counter-example: below modding-books 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

58
Zzhu_zong·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

430
Rrase·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

400
Rrase·28 minutes ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

423
Cchen_dev·2 days ago

This is not a modding-books problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

77
Kkernel_panic·just now

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

348
Ddev_zhouOP·2 days agoedited

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

251
KkiteOP·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

160
Cchen_dev·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

133
Cchen_dev·3 minutes ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

8
Bbob_chen·2 hours ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

428
Rrase·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

55
Aalice_devMod·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

37
Aalice_dev·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

2
Ttang_hao·2 days ago

I just read the modding-books source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

74
Cchen_dev·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

2
Zzhu_zong·28 minutes ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

333
Lli_ming·just now

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

117
Mmike_xu·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

15
Ddev_zhouMod·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

84
Cchen_dev·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

339
Kkernel_panic·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

333
Rran_bo·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

322
Cchen_devOP·2 days agoedited

This is not a modding-books problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

253
Nnikic·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

79
Zzhou_yi·5 hours ago

This is not a modding-books problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

300
Sslow_queryOP·just now

One counter-example: below modding-books 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

97
Sslow_queryOP·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

172
Rran_bo·2 days ago

This is not a modding-books problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

111
Bbob_chenMod·1 hour ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

122
Aalice_dev·2 days ago

One counter-example: below modding-books 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

429
Sslow_query·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

203
Nnikic·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

5
Ttang_haoMod·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

188
Hhuang_ke·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

273
Kkernel_panic·2 days agoedited

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

269
Ttang_hao·2 days agoedited

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

214
Ttang_haoOP·2 hours agoedited

Saved. I am reworking this area this week — this saves a lot of wrong turns.

146
Kkite·2 days ago

This is not a modding-books problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

94
Sslow_query·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

193
Kkernel_panic·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

58
Wwinter·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

179
Rrase·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

137
Sslow_query·12 minutes agoedited

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

180
Ddev_zhou·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

165
Wwinter·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

178
Nnikic·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

179
Wwinter·2 days ago

I just read the modding-books source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

218
Kkernel_panic·just now

Saved. I am reworking this area this week — this saves a lot of wrong turns.

1
Cchen_dev·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

205
Wwinter·2 days agoedited

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

174
Kkernel_panic·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

22
Zzhu_zong·1 hour ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

170
Nnikic·3 minutes ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

166
Oops_wang·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

37
Ttang_hao·2 days ago

One counter-example: below modding-books 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

66
Sswoole_leeOP·28 minutes agoedited

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

102
Sswoole_lee·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

465
Aalice_dev·just now

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

51
Rran_bo·1 hour ago

This is not a modding-books problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

485
Mmike_xu·3 minutes agoedited

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

7
Ddev_zhouMod·2 hours ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

269
Kkernel_panic·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

157
Sslow_query·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

119
Nnikic·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

10
Aalice_devOP·2 days ago

One counter-example: below modding-books 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

66
Bbob_chen·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

29
Sswoole_lee·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

111
Cchen_dev·2 days ago

This is not a modding-books problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

104
Sswoole_leeOP·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

40
RraseOP·yesterday

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

24
Rran_bo·2 days ago

I just read the modding-books source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

81
Rrase·2 days agoedited

I just read the modding-books source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

66
Wwinter·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

5
Wwinter·28 minutes ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

56
Lli_ming·2 days ago

I just read the modding-books source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

461
Sslow_query·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

251
Kkernel_panic·5 hours ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

51
Zzhou_yi·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

51
Nnikic·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

416
Wwinter·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

40
Ttang_haoOP·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

459
Kkite·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

254
Ttang_hao·2 days ago

This is not a modding-books problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

39
Sswoole_lee·just now

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

37
Nnikic·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

32
Sswoole_lee·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

24
Oops_wang·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

24
Cchen_dev·just now

I just read the modding-books source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

19
Zzhu_zong·2 days ago

I just read the modding-books source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

9
Aalice_dev·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

8
Kkernel_panic·just now

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

7
Rrase·1 hour ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

411
Zzhou_yi·3 minutes ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

479
Rrase·yesterday

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

374
Kkernel_panic·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

191
Rran_bo·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

4
Oops_wang·2 days agoLevel 6

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

310
Hhuang_ke·2 days agoLevel 6

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

205
Mmike_xu·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

1
Aalice_devMod·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

186
Rran_bo·2 days ago

One counter-example: below modding-books 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

6
Sslow_query·28 minutes ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

2
Aalice_dev·2 days ago

One counter-example: below modding-books 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

2
Mmike_xu·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

35
Wwinter·2 days agoedited

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

116
Mmike_xu·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

1
Sswoole_leeOP·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

431
Zzhou_yi·1 hour ago

I just read the modding-books source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

1
Oops_wangMod·2 hours ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

1
Rrase·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

9
Ttang_hao·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

1
Llinlin·2 days ago

One counter-example: below modding-books 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

3
Rran_bo·1 hour agoedited

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

1

This is the post detail page /en/c/modding-books/post/p2. Posts and comments are generated deterministically from a seeded PRNG, so the same post always renders the same content and the link can be shared, reloaded and indexed. In production this page reads MySQL for the post, Redis for hot-post caching, and fetches the whole comment tree in a single query on the path column.

See the database schema →