React · Architecture · After reading the react-design core source I finally understand the trade-offs

934
REr/react-design·posted by dev_zhou·yesterdayJobsLocked

After reading the react-design core source I finally understand the trade-offs

Most react-design articles stop at "how to use it" and never cover "when not to use it". This is an attempt at the second half.

One last trap: in container environments remember to adjust the memory-related parameters in step. Otherwise the host limit and the process expectation disagree, and the symptom is intermittent, unreproducible failure.

Image placeholder · object storage in production
POST /api/uploads → CDN origin pull
529 comments

529 comments

· first 120 loaded
M
Aalice_dev·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

503
Wwinter·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

103
Zzhu_zongMod·3 minutes ago

This is not a react-design problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

459
Kkernel_panic·12 minutes ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

2
Oops_wang·2 days agoedited

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

445
Oops_wang·yesterday

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

231
Hhuang_ke·2 days ago

One counter-example: below react-design 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

442
KkiteOP·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

369
Kkernel_panic·12 minutes ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

419
Ddev_zhou·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

477
Kkernel_panic·28 minutes ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

134
Zzhou_yi·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

484
Lli_mingOP·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

412
Sswoole_lee·2 days ago

This is not a react-design problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

102
Hhuang_ke·2 days agoLevel 6

I just read the react-design source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

498
Zzhu_zongOP·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

18
Ddev_zhou·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

315
Ddev_zhou·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

237
Aalice_dev·2 days agoLevel 6

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

113
Lli_ming·2 days agoLevel 6

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

11
Hhuang_ke·5 hours ago

I just read the react-design source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

79
Kkite·2 days agoLevel 6

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

506
Wwinter·2 days agoLevel 6

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

393
Sswoole_lee·2 days agoLevel 6

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

49
Zzhou_yi·2 days agoedited

One counter-example: below react-design 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

11
Zzhu_zong·2 days agoedited

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

19
KkiteOP·2 days agoLevel 6

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

7
Sswoole_leeMod·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

3
Lli_mingOP·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

365
Lli_ming·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

212
NnikicOP·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

1
Zzhu_zong·2 days ago

This is not a react-design problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

138
Rrase·2 days agoLevel 6

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

447
Wwinter·28 minutes ago

I just read the react-design source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

136
Mmike_xu·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

259
Hhuang_ke·2 days ago

One counter-example: below react-design 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

48
Kkite·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

400
RraseOP·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

328
Ddev_zhou·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

376
Mmike_xu·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

2
Zzhu_zong·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

324
Rran_bo·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

321
Llinlin·2 days agoedited

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

312
Bbob_chen·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

1
Mmike_xu·2 days ago

I just read the react-design source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

471
Cchen_dev·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

516
Rran_boOPMod·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

201
Cchen_dev·2 days ago

This is not a react-design problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

423
Wwinter·2 days agoedited

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

40
Llinlin·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

36
KkiteOP·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

204
Llinlin·2 days ago

One counter-example: below react-design 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

256
Kkite·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

290
Cchen_dev·2 days ago

I just read the react-design source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

50
Mmike_xu·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

236
Llinlin·12 minutes ago

This is not a react-design problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

222
Llinlin·5 hours ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

433
NnikicMod·2 hours ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

248
Cchen_dev·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

16
Bbob_chen·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

1
Ttang_haoMod·yesterday

One counter-example: below react-design 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

28
Nnikic·2 days ago

One counter-example: below react-design 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

218
Mmike_xu·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

211
Lli_ming·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

205
Kkernel_panic·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

202
Aalice_dev·2 hours ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

200
Llinlin·2 days ago

I just read the react-design source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

142
Zzhou_yi·2 days ago

One counter-example: below react-design 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

184
Cchen_dev·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

158
Lli_ming·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

4
Sslow_query·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

449
Rran_bo·2 days ago

I just read the react-design source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

154
Oops_wang·2 days ago

This is not a react-design problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

148
Kkite·12 minutes agoedited

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

147
Lli_ming·2 days agoedited

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

122
Ttang_hao·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

24
WwinterOP·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

384
Kkernel_panic·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

12
Rran_bo·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

185
Cchen_dev·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

174
Lli_mingMod·1 hour ago

This is not a react-design problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

33
Rran_bo·2 days agoedited

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

315
Aalice_dev·12 minutes ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

7
Hhuang_ke·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

1
Lli_ming·1 hour agoLevel 6

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

30
Oops_wang·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

137
Kkite·12 minutes ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

135
RraseMod·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

443
Oops_wang·yesterday

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

248
Llinlin·just now

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

129
Llinlin·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

128
Oops_wang·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

107
Rran_bo·3 minutes ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

79
Ddev_zhou·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

428
Zzhu_zong·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

41
Rran_bo·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

8
NnikicMod·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

2
Kkernel_panic·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

73
Llinlin·2 days agoedited

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

67
Ttang_hao·3 minutes ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

65
Cchen_dev·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

7
Cchen_dev·1 hour ago

I just read the react-design source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

65
Cchen_devMod·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

1
Sslow_queryOP·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

1
Sslow_query·2 days agoedited

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

42
Mmike_xu·1 hour agoedited

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

469
Lli_ming·28 minutes ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

24
Ttang_haoOP·2 days ago

One counter-example: below react-design 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

445
Rran_bo·3 minutes ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

22
Zzhu_zong·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

20
Rran_boOP·2 days agoedited

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

3
Rrase·2 days ago

This is not a react-design problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

183
Oops_wangMod·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

1
Kkite·2 hours ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

1
Kkernel_panic·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

15
Mmike_xu·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

7
Llinlin·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

4
Aalice_dev·just now

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

3
Bbob_chen·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

3
Bbob_chen·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

3

This is the post detail page /en/c/react-design/post/p1. Posts and comments are generated deterministically from a seeded PRNG, so the same post always renders the same content and the link can be shared, reloaded and indexed. In production this page reads MySQL for the post, Redis for hot-post caching, and fetches the whole comment tree in a single query on the path column.

See the database schema →