Clojure · Job board · Built a visual debugging panel for clojure-jobs over the weekend

1.6K
CLr/clojure-jobs·posted by bob_chen·2 days agoTutorial

Built a visual debugging panel for clojure-jobs over the weekend

Most clojure-jobs articles stop at "how to use it" and never cover "when not to use it". This is an attempt at the second half.

On trade-offs, my view is this: if nobody on the team owns this area long-term, do not introduce a second mechanism. With two coexistence you first have to work out which one is even in play when things break, and that costs far more than the performance you saved.

Self-host — full control44%
Managed service — less work28%
Hybrid: core self-hosted17%
Not decided yet11%

2678 votes total

419 comments

419 comments

· first 120 loaded
M
Cchen_dev·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

520
Llinlin·28 minutes ago

I just read the clojure-jobs source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

494
Aalice_dev·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

471
Rran_boOP·just now

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

399
Nnikic·2 days ago

One counter-example: below clojure-jobs 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

457
Ttang_hao·2 days ago

One counter-example: below clojure-jobs 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

432
Kkite·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

415
Sslow_query·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

407
Sslow_query·12 minutes ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

390
Nnikic·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

368
Rran_bo·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

230
Mmike_xu·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

305
Lli_ming·3 minutes ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

252
Oops_wangOP·12 minutes agoedited

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

186
Mmike_xu·12 minutes ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

322
Hhuang_keOP·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

12
Wwinter·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

118
Aalice_dev·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

35
Bbob_chenMod·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

15
Sslow_query·12 minutes ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

138
Ttang_hao·1 hour ago

This is not a clojure-jobs problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

101
Sswoole_leeMod·yesterday

I just read the clojure-jobs source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

25
Ddev_zhou·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

8
Nnikic·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

224
Ttang_hao·2 days ago

One counter-example: below clojure-jobs 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

8
Oops_wang·2 days agoedited

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

102
Wwinter·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

221
Kkernel_panic·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

206
Sswoole_lee·1 hour ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

332
Mmike_xu·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

474
Aalice_dev·28 minutes ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

8
Sswoole_lee·28 minutes ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

166
Bbob_chen·2 days agoedited

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

69
Kkernel_panic·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

16
Aalice_dev·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

187
Sslow_query·2 days ago

One counter-example: below clojure-jobs 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

158
NnikicOP·28 minutes ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

127
Lli_ming·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

141
Rran_bo·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

80
Wwinter·1 hour agoedited

Saved. I am reworking this area this week — this saves a lot of wrong turns.

129
Sswoole_lee·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

121
Oops_wang·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

109
Mmike_xu·2 days ago

This is not a clojure-jobs problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

109
Mmike_xu·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

107
Aalice_dev·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

295
Mmike_xu·5 hours ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

61
Bbob_chen·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

14
Rran_bo·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

53
Sslow_query·2 hours ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

39
Cchen_dev·12 minutes ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

378
Sswoole_lee·2 days ago

I just read the clojure-jobs source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

414
Nnikic·yesterday

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

363
Cchen_devMod·2 days ago

One counter-example: below clojure-jobs 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

306
Sswoole_lee·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

13
Zzhou_yi·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

7
Oops_wangOP·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

8
Kkite·2 hours ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

1
Cchen_devOPMod·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

338
Wwinter·2 days ago

This is not a clojure-jobs problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

84
Rran_bo·2 days agoedited

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

95
Oops_wang·1 hour ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

37
Lli_ming·2 days ago

I just read the clojure-jobs source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

37
Cchen_dev·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

36
Sslow_query·1 hour ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

3
Zzhou_yi·2 days ago

I just read the clojure-jobs source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

163
Sslow_query·2 days agoedited

I just read the clojure-jobs source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

27
Kkernel_panic·yesterday

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

22
Zzhu_zong·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

374
Bbob_chen·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

122
Sslow_query·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

233
Bbob_chen·2 days agoedited

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

21
Sslow_query·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

18
Nnikic·2 days agoedited

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

16
Bbob_chen·yesterday

Saved. I am reworking this area this week — this saves a lot of wrong turns.

15
Ddev_zhou·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

12
Bbob_chen·just now

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

10
Aalice_devMod·2 hours ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

6
Ttang_hao·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

247
Kkernel_panic·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

60
Zzhou_yiOP·28 minutes ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

357
Kkite·12 minutes ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

363
Sswoole_leeOP·yesterdayLevel 6

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

396
Aalice_devOP·2 days agoLevel 6

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

325
Sswoole_lee·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

279
Oops_wang·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

65
Sswoole_lee·2 days agoLevel 6

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

457
Oops_wang·yesterday

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

312
Rran_bo·2 days ago

This is not a clojure-jobs problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

24
Oops_wang·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

366
Kkite·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

249
Hhuang_ke·2 hours ago

I just read the clojure-jobs source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

36
Hhuang_ke·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

508
Lli_ming·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

499
Lli_ming·2 days agoLevel 6

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

18
Lli_mingOP·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

334
Lli_ming·2 days ago

This is not a clojure-jobs problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

174
Kkernel_panic·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

2
NnikicMod·2 days agoLevel 6

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

377
Lli_mingMod·2 days agoLevel 6

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

278
Ddev_zhou·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

10
Sswoole_lee·12 minutes ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

8
Mmike_xu·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

56
Bbob_chen·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

128
Aalice_dev·2 days ago

One counter-example: below clojure-jobs 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

201
Kkite·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

121
Rran_boOP·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

264
Cchen_dev·2 hours agoedited

This is not a clojure-jobs problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

14
Zzhou_yi·3 minutes ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

445
Sswoole_lee·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

113
Sswoole_lee·2 days agoedited

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

383
Rran_bo·yesterday

One counter-example: below clojure-jobs 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

82
Zzhu_zong·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

139
Bbob_chen·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

7
Sslow_query·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

254
Mmike_xuMod·5 hours ago

This is not a clojure-jobs problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

122
Rran_bo·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

2
Mmike_xu·2 days agoedited

I just read the clojure-jobs source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

2
Cchen_dev·1 hour ago

One counter-example: below clojure-jobs 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

1
Zzhou_yi·5 hours ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

36
Zzhu_zong·2 days ago

This is not a clojure-jobs problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

1

This is the post detail page /en/c/clojure-jobs/post/p2. Posts and comments are generated deterministically from a seeded PRNG, so the same post always renders the same content and the link can be shared, reloaded and indexed. In production this page reads MySQL for the post, Redis for hot-post caching, and fetches the whole comment tree in a single query on the path column.

See the database schema →