OWASP · Open source · Those easily-missed type details in owasp-opensource

518
OWr/owasp-opensource·posted by bob_chen·3 days agoReview

Those easily-missed type details in owasp-opensource

Some background first. Our setup is owasp-opensource plus three downstream services, seven figures of daily requests, peaking around nine in the evening.

One last trap: in container environments remember to adjust the memory-related parameters in step. Otherwise the host limit and the process expectation disagree, and the symptom is intermittent, unreproducible failure.

Image placeholder · object storage in production
POST /api/uploads → CDN origin pull
129 comments

129 comments

· first 120 loaded
M
Wwinter·2 days ago

I just read the owasp-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

483
Sswoole_lee·1 hour ago

I just read the owasp-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

479
LlinlinOP·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

129
Rran_bo·1 hour ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

370
Zzhu_zong·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

326
Nnikic·just nowedited

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

270
Mmike_xu·1 hour ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

213
Kkernel_panic·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

279
Hhuang_ke·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

119
Zzhou_yi·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

257
Sslow_query·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

152
Rran_bo·2 days ago

I just read the owasp-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

136
Ttang_hao·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

140
Aalice_dev·2 days agoLevel 6

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

133
Ttang_hao·2 days agoLevel 6

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

63
Sslow_query·1 hour ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

103
Nnikic·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

392
Zzhou_yi·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

334
Ddev_zhou·2 days ago

I just read the owasp-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

322
Hhuang_ke·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

318
Ttang_hao·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

253
Rran_bo·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

1
Rran_bo·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

234
Ddev_zhou·yesterday

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

2
Sswoole_leeOP·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

246
Kkernel_panic·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

1
Kkernel_panic·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

216
Bbob_chen·2 days agoedited

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

1
Wwinter·2 days agoedited

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

221
Zzhou_yi·2 days agoedited

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

81
Ttang_hao·2 days ago

I just read the owasp-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

478
Aalice_devOP·12 minutes agoedited

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

149
Rran_bo·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

197
Sswoole_lee·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

181
Rran_bo·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

146
Llinlin·2 days agoedited

Saved. I am reworking this area this week — this saves a lot of wrong turns.

115
Llinlin·just now

This is not a owasp-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

150
Hhuang_ke·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

141
Zzhou_yi·just now

One counter-example: below owasp-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

140
Cchen_dev·2 days agoedited

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

6
Aalice_devOP·2 days ago

This is not a owasp-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

74
Rrase·2 days agoedited

One counter-example: below owasp-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

123
Kkernel_panic·12 minutes ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

111
Ttang_haoOP·2 days ago

One counter-example: below owasp-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

51
Sslow_query·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

108
Zzhou_yi·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

252
Aalice_dev·2 days ago

This is not a owasp-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

28
Cchen_dev·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

83
Aalice_dev·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

138
Ddev_zhouMod·2 days agoedited

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

41
Rran_bo·2 days ago

I just read the owasp-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

107
NnikicOP·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

46
Aalice_dev·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

98
KkiteOP·2 days agoedited

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

33
Mmike_xu·3 minutes ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

88
Kkernel_panicOP·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

62
Ttang_haoMod·5 hours ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

85
Ttang_hao·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

277
Ttang_hao·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

77
Hhuang_ke·2 days ago

This is not a owasp-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

80
Sslow_query·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

78
KkiteMod·yesterday

I just read the owasp-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

1
Ddev_zhouOP·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

12
Zzhou_yi·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

69
Ttang_hao·2 days agoedited

One counter-example: below owasp-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

65
Lli_ming·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

36
Nnikic·2 days ago

This is not a owasp-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

29
Ttang_hao·2 days agoedited

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

14
Ddev_zhou·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

283
Lli_ming·just now

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

4
Cchen_dev·2 days agoedited

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

466
Bbob_chen·12 minutes ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

252
Zzhu_zong·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

2
Zzhou_yi·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

507
Hhuang_ke·12 minutes ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

1
Rran_bo·3 minutes ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

219
Ttang_hao·12 minutes ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

458
Wwinter·5 hours ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

395
Aalice_dev·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

437
Kkite·2 days agoLevel 6

This is not a owasp-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

83
Hhuang_ke·2 days agoLevel 6

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

17
Mmike_xuMod·2 hours agoLevel 6

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

6
Mmike_xu·1 hour ago

This is not a owasp-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

186
Zzhou_yi·yesterdayLevel 6

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

31
Mmike_xu·2 days agoLevel 6

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

10
Kkite·5 hours agoeditedLevel 6

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

1
Ttang_hao·2 days ago

One counter-example: below owasp-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

167
Zzhou_yi·2 days agoLevel 6

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

457
Kkernel_panic·1 hour agoedited

Saved. I am reworking this area this week — this saves a lot of wrong turns.

16
Hhuang_ke·12 minutes ago

One counter-example: below owasp-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

26
Cchen_devOP·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

54
Cchen_dev·2 days agoedited

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

6
Mmike_xu·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

102
Oops_wang·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

275
Llinlin·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

169
Nnikic·1 hour ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

116
Aalice_dev·2 days agoLevel 6

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

21
Cchen_dev·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

34
Sswoole_lee·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

33
NnikicOP·2 days agoLevel 6

I just read the owasp-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

514
Rran_bo·2 days ago

One counter-example: below owasp-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

45
Mmike_xu·2 days agoedited

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

145
Sswoole_lee·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

1
Ddev_zhou·5 hours ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

31
Kkernel_panic·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

2
Sslow_query·just now

Saved. I am reworking this area this week — this saves a lot of wrong turns.

11
Rran_bo·12 minutes ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

484
WwinterMod·2 days agoLevel 6

One counter-example: below owasp-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

284
Rrase·28 minutes ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

476
Hhuang_ke·2 days agoLevel 6

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

246
Lli_ming·2 days agoeditedLevel 6

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

128
Sslow_queryOP·2 days agoLevel 6

Saved. I am reworking this area this week — this saves a lot of wrong turns.

1
Rran_boOP·2 days agoedited

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

155
Zzhou_yiOP·2 days agoLevel 6

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

147
NnikicOP·2 days agoLevel 6

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

14
Kkite·3 minutes ago

This is not a owasp-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

3
Ddev_zhou·2 hours ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

78
Mmike_xuOP·2 days agoedited

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

72
Aalice_dev·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

1
Nnikic·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

1

This is the post detail page /en/c/owasp-opensource/post/p3. Posts and comments are generated deterministically from a seeded PRNG, so the same post always renders the same content and the link can be shared, reloaded and indexed. In production this page reads MySQL for the post, Redis for hot-post caching, and fetches the whole comment tree in a single query on the path column.

See the database schema →