OWASP · Open source · The edge cases the owasp-opensource docs never spell out

2.4K
OWr/owasp-opensource·posted by kernel_panic·3 hours agoReview

The edge cases the owasp-opensource docs never spell out

It took me two weeks of on-and-off digging and plenty of wrong turns. Writing the process down as it happened so the next person spends less time.

What genuinely surprised me was the tail. The average looked great while P99 jumped by an order of magnitude past some threshold. The cause was not owasp-opensource itself but our upstream connection reuse — the load test traffic was too clean and hid the long-tail requests.

One last trap: in container environments remember to adjust the memory-related parameters in step. Otherwise the host limit and the process expectation disagree, and the symptom is intermittent, unreproducible failure.

Sstackoverflow.comExternal link · opens in a new tab
1080 comments

1080 comments

· first 120 loaded
M
KkiteMod·yesterday

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

474
Lli_mingMod·2 days ago

This is not a owasp-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

15
Oops_wangOP·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

1
Mmike_xu·yesterday

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

71
Mmike_xuMod·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

316
Ddev_zhou·5 hours ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

3
Cchen_dev·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

466
Aalice_dev·just now

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

463
Llinlin·2 days ago

One counter-example: below owasp-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

434
Ddev_zhou·3 minutes ago

This is not a owasp-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

422
Zzhu_zong·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

385
Ttang_hao·3 minutes ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

360
Bbob_chenOP·3 minutes ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

322
Zzhou_yi·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

460
RraseOP·2 days agoedited

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

27
Llinlin·yesterday

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

28
Zzhu_zong·2 days agoedited

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

364
Sslow_query·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

515
Aalice_dev·just now

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

322
Lli_ming·just nowLevel 6

Saved. I am reworking this area this week — this saves a lot of wrong turns.

356
Kkernel_panic·2 days agoLevel 6

Saved. I am reworking this area this week — this saves a lot of wrong turns.

260
Kkite·2 days agoLevel 6

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

54
Oops_wang·1 hour ago

This is not a owasp-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

20
Nnikic·3 minutes agoLevel 6

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

18
Zzhu_zong·1 hour ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

199
Kkite·yesterday

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

388
Rrase·yesterdayLevel 6

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

253
Rran_bo·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

9
Rrase·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

208
Rrase·2 days agoeditedLevel 6

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

288
RraseMod·2 days agoeditedLevel 6

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

247
Wwinter·2 days agoLevel 6

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

197
Aalice_dev·2 days agoedited

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

25
Mmike_xu·2 days ago

This is not a owasp-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

358
Sswoole_lee·just now

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

355
Bbob_chenOP·12 minutes ago

I just read the owasp-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

337
Sslow_query·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

178
Aalice_dev·1 hour agoedited

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

87
Sslow_query·yesterday

This is not a owasp-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

428
Zzhu_zong·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

208
Hhuang_ke·12 minutes ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

1
Cchen_dev·yesterday

One counter-example: below owasp-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

165
Kkite·5 hours ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

519
Cchen_dev·12 minutes ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

112
Ttang_hao·just now

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

41
Wwinter·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

458
Ttang_hao·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

92
Zzhu_zong·2 days ago

I just read the owasp-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

333
Rrase·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

331
Rran_bo·2 hours ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

372
Kkite·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

309
Cchen_dev·yesterday

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

302
Hhuang_ke·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

286
Zzhou_yiOP·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

217
Cchen_dev·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

4
Llinlin·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

489
Hhuang_keOP·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

8
Sswoole_lee·2 hours ago

One counter-example: below owasp-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

164
Sslow_query·2 hours ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

362
Zzhou_yi·5 hours ago

I just read the owasp-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

160
Zzhu_zong·2 days ago

I just read the owasp-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

140
Kkernel_panic·1 hour ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

132
Rran_bo·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

109
Bbob_chenMod·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

107
Lli_ming·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

104
Cchen_dev·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

38
Nnikic·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

96
Sslow_queryOP·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

32
Zzhu_zongOP·just now

One counter-example: below owasp-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

134
Hhuang_ke·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

311
Zzhou_yiMod·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

201
Kkite·2 days agoedited

One counter-example: below owasp-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

10
Cchen_devOP·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

14
Zzhou_yi·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

73
Rran_boOP·just now

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

3
Nnikic·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

59
Oops_wang·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

452
Rrase·2 days ago

I just read the owasp-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

56
Kkernel_panic·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

54
Wwinter·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

48
Cchen_dev·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

37
Sswoole_lee·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

35
Ttang_haoOP·2 days ago

This is not a owasp-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

405
Zzhu_zong·2 days ago

This is not a owasp-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

504
Kkernel_panic·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

1
Cchen_dev·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

147
Rran_boMod·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

29
Sslow_query·2 days ago

This is not a owasp-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

26
Hhuang_ke·2 days agoedited

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

24
Kkernel_panic·12 minutes ago

One counter-example: below owasp-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

23
Oops_wang·just now

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

16
Zzhou_yiMod·just now

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

110
Sslow_query·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

256
Hhuang_ke·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

88
Wwinter·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

10
Mmike_xu·1 hour ago

One counter-example: below owasp-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

15
Zzhou_yi·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

14
Sswoole_lee·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

10
Zzhou_yi·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

8
Llinlin·just now

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

359
Zzhou_yi·2 days agoedited

I just read the owasp-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

1
Zzhou_yi·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

340
Mmike_xu·2 days ago

One counter-example: below owasp-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

5
Rran_boOP·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

390
Bbob_chen·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

29
Bbob_chen·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

224
Hhuang_ke·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

267
Cchen_dev·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

510
Nnikic·1 hour ago

I just read the owasp-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

162
Bbob_chen·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

471
Cchen_dev·2 days agoedited

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

43
Zzhou_yi·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

11
Sslow_query·2 days ago

This is not a owasp-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

3
Kkernel_panic·12 minutes ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

2
Zzhou_yi·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

499
Kkernel_panic·2 days agoedited

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

2
KkiteMod·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

2
Sswoole_lee·2 days ago

I just read the owasp-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

18
Oops_wang·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

1
Ttang_hao·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

1

This is the post detail page /en/c/owasp-opensource/post/p8. Posts and comments are generated deterministically from a seeded PRNG, so the same post always renders the same content and the link can be shared, reloaded and indexed. In production this page reads MySQL for the post, Redis for hot-post caching, and fetches the whole comment tree in a single query on the path column.

See the database schema →