Clojure · Job board · The clojure-jobs execution flow in one diagram (with sequence chart)

532
CLr/clojure-jobs·posted by li_ming·3 hours agoPostmortemLocked

The clojure-jobs execution flow in one diagram (with sequence chart)

Some background first. Our setup is clojure-jobs plus three downstream services, seven figures of daily requests, peaking around nine in the evening.

On trade-offs, my view is this: if nobody on the team owns this area long-term, do not introduce a second mechanism. With two coexistence you first have to work out which one is even in play when things break, and that costs far more than the performance you saved.

Documentation first44%
Source code first28%
Just ask someone17%
Run a demo and learn by error11%

904 votes total

165 comments

165 comments

· first 120 loaded
M
Rran_bo·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

509
Rran_boMod·2 days agoedited

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

502
WwinterMod·2 hours ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

59
Lli_ming·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

489
Ddev_zhouMod·2 days agoedited

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

14
Ttang_hao·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

464
Kkernel_panic·2 days agoedited

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

247
Nnikic·2 days ago

I just read the clojure-jobs source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

3
Bbob_chen·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

53
Mmike_xu·5 hours agoedited

Saved. I am reworking this area this week — this saves a lot of wrong turns.

434
Lli_ming·3 minutes ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

390
Zzhou_yi·2 days ago

I just read the clojure-jobs source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

368
Ddev_zhou·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

222
Lli_ming·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

370
Sswoole_lee·12 minutes ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

27
Wwinter·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

34
Hhuang_ke·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

363
Ttang_hao·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

360
Sswoole_lee·2 days ago

I just read the clojure-jobs source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

353
KkiteOP·2 days agoedited

One counter-example: below clojure-jobs 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

266
Oops_wang·just now

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

396
Llinlin·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

154
Bbob_chenOP·2 hours ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

262
Llinlin·yesterday

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

322
Kkite·2 days agoedited

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

321
Lli_ming·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

315
Rrase·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

304
Ttang_hao·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

186
Bbob_chen·2 days ago

One counter-example: below clojure-jobs 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

10
Zzhou_yi·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

398
Rran_bo·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

278
Kkite·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

248
Oops_wang·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

248
Bbob_chen·2 days ago

This is not a clojure-jobs problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

238
Hhuang_ke·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

5
Aalice_dev·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

230
Rrase·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

228
Rrase·2 days ago

This is not a clojure-jobs problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

218
Zzhou_yi·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

210
Ddev_zhou·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

194
Ttang_haoMod·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

182
Oops_wang·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

165
Sswoole_lee·1 hour agoedited

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

125
Kkite·2 days ago

I just read the clojure-jobs source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

123
Sswoole_lee·12 minutes ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

104
Zzhu_zong·12 minutes ago

One counter-example: below clojure-jobs 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

103
Zzhou_yi·1 hour ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

499
Ttang_hao·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

193
Nnikic·2 days ago

This is not a clojure-jobs problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

421
Rran_bo·just now

I just read the clojure-jobs source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

37
Kkite·2 days ago

I just read the clojure-jobs source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

3
Lli_ming·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

56
Kkernel_panic·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

8
Ttang_hao·12 minutes ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

163
Llinlin·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

325
Rran_bo·yesterday

Saved. I am reworking this area this week — this saves a lot of wrong turns.

408
Kkernel_panic·12 minutes ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

314
Nnikic·2 days agoedited

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

77
Rran_bo·2 days agoLevel 6

One counter-example: below clojure-jobs 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

65
Sslow_query·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

34
Kkernel_panic·2 days agoLevel 6

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

18
Wwinter·2 days agoedited

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

202
Hhuang_ke·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

1
Aalice_dev·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

25
Hhuang_ke·2 days ago

This is not a clojure-jobs problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

2
Zzhou_yi·2 days agoeditedLevel 6

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

3
Zzhou_yiOP·2 hours agoedited

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

75
Oops_wang·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

401
Ttang_hao·2 days agoedited

This is not a clojure-jobs problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

177
Ddev_zhou·2 days ago

I just read the clojure-jobs source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

51
Kkite·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

434
Cchen_dev·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

7
Rran_bo·2 days agoedited

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

132
Ttang_hao·2 days agoedited

This is not a clojure-jobs problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

76
Llinlin·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

59
Aalice_dev·yesterdayedited

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

47
Zzhou_yiOP·yesterday

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

41
Bbob_chenOP·2 days ago

One counter-example: below clojure-jobs 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

8
Rran_bo·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

516
Rran_bo·1 hour ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

330
Nnikic·2 hours ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

3
Rran_bo·2 days agoLevel 6

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

164
Mmike_xu·2 hours ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

1
Lli_mingOP·1 hour ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

9
Llinlin·5 hours ago

One counter-example: below clojure-jobs 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

23
Llinlin·3 minutes ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

366
Sslow_query·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

330
Aalice_devMod·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

2
Ddev_zhou·2 hours ago

This is not a clojure-jobs problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

1
Llinlin·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

311
Kkernel_panicOP·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

118
Oops_wang·2 days agoLevel 6

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

13
Ddev_zhou·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

42
Nnikic·2 days agoedited

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

40
Hhuang_keMod·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

33
Ddev_zhou·28 minutes ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

32
Sslow_query·1 hour ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

30
Kkite·3 minutes ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

38
Cchen_dev·2 days agoedited

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

420
Mmike_xu·2 days agoedited

Saved. I am reworking this area this week — this saves a lot of wrong turns.

423
Aalice_devOP·2 days agoedited

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

16
Rrase·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

5
Oops_wang·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

210
Kkite·28 minutes ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

28
Cchen_devMod·2 hours ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

24
Rrase·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

23
Mmike_xu·2 days ago

I just read the clojure-jobs source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

52
Cchen_dev·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

18
Sswoole_lee·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

10
Nnikic·2 days agoedited

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

6
Bbob_chen·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

4
Rran_bo·2 days ago

One counter-example: below clojure-jobs 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

2
Kkite·12 minutes ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

1
Wwinter·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

284
Bbob_chen·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

31
Aalice_dev·2 days ago

This is not a clojure-jobs problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

256
Oops_wang·2 days ago

One counter-example: below clojure-jobs 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

1
Lli_ming·1 hour ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

1
Wwinter·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

1
Mmike_xu·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

1

This is the post detail page /en/c/clojure-jobs/post/p8. Posts and comments are generated deterministically from a seeded PRNG, so the same post always renders the same content and the link can be shared, reloaded and indexed. In production this page reads MySQL for the post, Redis for hot-post caching, and fetches the whole comment tree in a single query on the path column.

See the database schema →