Linux Community · Two weeks with linux: what feels right and what drives me up the wall

406
LIr/linux·posted by ops_wang·3 days agoPostmortem

Two weeks with linux: what feels right and what drives me up the wall

It took me two weeks of on-and-off digging and plenty of wrong turns. Writing the process down as it happened so the next person spends less time.

We also fixed monitoring along the way: replaced average-based alerts with percentiles and split them per endpoint. False alerts dropped by about seventy percent and the on-call rotation visibly cheered up.

Image placeholder · object storage in production
POST /api/uploads → CDN origin pull
227 comments

227 comments

· first 120 loaded
M
WwinterOP·2 days agoedited

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

370
Mmike_xu·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

43
Aalice_dev·1 hour ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

378
Bbob_chen·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

506
Cchen_dev·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

311
Oops_wang·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

266
Ttang_hao·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

288
Cchen_dev·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

104
Wwinter·2 days ago

This is not a linux problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

9
Aalice_devOP·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

203
Wwinter·2 days ago

I just read the linux source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

249
Sswoole_lee·2 days ago

One counter-example: below linux 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

219
Rrase·2 days agoedited

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

214
Oops_wangOP·12 minutes ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

147
Sswoole_lee·1 hour ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

263
Sswoole_leeOP·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

2
Lli_ming·12 minutes ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

317
Lli_ming·2 days agoedited

One counter-example: below linux 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

205
Hhuang_ke·2 days agoedited

I just read the linux source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

199
Oops_wang·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

198
Kkite·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

447
Bbob_chen·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

38
Llinlin·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

438
Kkernel_panic·2 days ago

This is not a linux problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

189
Wwinter·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

1
Sslow_query·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

180
Cchen_dev·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

174
Sslow_query·2 hours agoedited

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

156
Aalice_dev·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

154
Mmike_xu·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

133
WwinterOP·2 hours ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

64
Mmike_xu·12 minutes ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

216
Wwinter·1 hour ago

This is not a linux problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

157
Lli_ming·just now

This is not a linux problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

131
Ttang_hao·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

4
Wwinter·2 days ago

One counter-example: below linux 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

117
KkiteMod·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

110
Nnikic·1 hour ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

102
Lli_mingMod·2 days ago

I just read the linux source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

95
Ddev_zhou·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

87
Kkernel_panic·2 days agoedited

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

86
Sswoole_lee·1 hour ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

83
LlinlinOP·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

100
Hhuang_keMod·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

37
Oops_wang·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

16
RraseMod·28 minutes agoedited

I just read the linux source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

3
Ttang_hao·2 days agoedited

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

81
Rran_bo·2 days agoedited

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

305
Zzhu_zong·3 minutes ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

1
Oops_wang·2 days agoedited

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

81
Kkite·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

79
Bbob_chen·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

77
Ddev_zhou·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

71
Bbob_chen·2 days agoedited

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

477
Lli_mingOP·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

11
Ddev_zhou·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

71
Cchen_dev·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

439
Rran_bo·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

201
Wwinter·5 hours ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

68
Zzhou_yiOP·2 hours ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

7
Kkite·12 minutes ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

155
Lli_ming·1 hour ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

274
Cchen_dev·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

219
Sswoole_lee·5 hours ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

481
Lli_ming·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

457
Lli_ming·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

3
Mmike_xu·just now

This is not a linux problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

83
Kkite·2 days ago

I just read the linux source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

466
Mmike_xu·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

25
Bbob_chenOP·2 days agoeditedLevel 6

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

251
Bbob_chen·2 days agoLevel 6

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

19
Sswoole_leeMod·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

129
Rran_bo·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

180
Sslow_query·3 minutes ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

49
Kkite·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

2
Rran_bo·2 days agoLevel 6

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

399
Ttang_hao·2 days agoLevel 6

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

292
Hhuang_keOP·2 days agoLevel 6

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

44
Wwinter·yesterday

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

36
Zzhou_yi·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

38
Zzhou_yi·2 days agoedited

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

1
Aalice_dev·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

123
Hhuang_ke·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

106
Nnikic·2 days ago

I just read the linux source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

131
Kkernel_panic·12 minutes ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

73
Rran_bo·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

63
Nnikic·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

337
Aalice_dev·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

5
Cchen_dev·2 hours ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

66
Hhuang_ke·2 days ago

This is not a linux problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

456
Rran_bo·2 days ago

One counter-example: below linux 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

9
Oops_wang·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

48
Rran_bo·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

2
Mmike_xu·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

44
Mmike_xu·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

41
Ttang_haoMod·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

40
Oops_wang·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

37
Oops_wang·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

133
Oops_wang·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

330
Sslow_query·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

234
Hhuang_ke·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

143
Llinlin·3 minutes ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

185
Bbob_chen·2 days agoedited

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

64
Lli_ming·1 hour ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

20
Cchen_dev·2 days ago

This is not a linux problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

372
Mmike_xu·yesterday

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

7
Sslow_query·2 days ago

One counter-example: below linux 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

13
Zzhou_yi·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

4
Nnikic·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

4
Kkite·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

46
Ddev_zhou·2 days ago

This is not a linux problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

3
Aalice_dev·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

3
Zzhou_yi·just now

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

2
Aalice_dev·just now

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

1
Cchen_dev·yesterday

One counter-example: below linux 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

1
Hhuang_ke·yesterday

I just read the linux source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

444
Lli_ming·2 days ago

One counter-example: below linux 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

1
Kkite·2 days agoedited

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

332
Sswoole_lee·2 days ago

One counter-example: below linux 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

1
Sswoole_lee·2 days ago

I just read the linux source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

1

This is the post detail page /en/c/linux/post/p3. Posts and comments are generated deterministically from a seeded PRNG, so the same post always renders the same content and the link can be shared, reloaded and indexed. In production this page reads MySQL for the post, Redis for hot-post caching, and fetches the whole comment tree in a single query on the path column.

See the database schema →