Ubuntu Community · I upgraded our production Ubuntu and hit these 11 landmines

771
Ubr/ubuntu·posted by huang_ke·3 days agoAnnouncement

I upgraded our production Ubuntu and hit these 11 landmines

It took me two weeks of on-and-off digging and plenty of wrong turns. Writing the process down as it happened so the next person spends less time.

Order of investigation, by return on effort: 1. Check downstream latency first — usually it is not your problem 2. Then pool hit rate and wait-queue length 3. Only then GC and allocation 4. Suspect the framework last

We also fixed monitoring along the way: replaced average-based alerts with percentiles and split them per endpoint. False alerts dropped by about seventy percent and the on-call rotation visibly cheered up.

One last trap: in container environments remember to adjust the memory-related parameters in step. Otherwise the host limit and the process expectation disagree, and the symptom is intermittent, unreproducible failure.

Worth noting: the official docs do cover this, just in a very inconspicuous spot. I only found it reading the source comments, where the author explains the reasoning — roughly "so that it degrades into predictable behaviour in extreme cases".

219 comments

219 comments

· first 120 loaded
M
Sswoole_leeMod·1 hour ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

503
Zzhu_zong·28 minutes ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

8
Hhuang_ke·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

417
Cchen_dev·just now

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

100
Zzhou_yi·1 hour ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

310
RraseMod·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

470
Zzhu_zong·2 days ago

One counter-example: below Ubuntu 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

1
Hhuang_ke·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

32
Llinlin·2 days agoedited

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

500
Zzhu_zong·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

9
Sslow_query·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

460
Lli_mingMod·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

386
Kkernel_panic·2 days agoedited

This is not a Ubuntu problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

352
Sslow_query·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

301
Kkernel_panic·2 days agoedited

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

347
Oops_wang·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

347
Ttang_hao·just now

I just read the Ubuntu source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

331
Sswoole_lee·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

297
Ddev_zhou·2 hours ago

This is not a Ubuntu problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

292
Sslow_queryOP·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

91
Llinlin·2 days agoedited

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

255
LlinlinOP·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

171
Cchen_dev·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

209
Cchen_dev·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

255
Sswoole_lee·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

1
Rrase·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

3
Sswoole_lee·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

124
Llinlin·2 days ago

This is not a Ubuntu problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

121
KkiteOP·5 hours agoedited

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

1
Sswoole_lee·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

289
Oops_wang·2 days ago

One counter-example: below Ubuntu 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

92
Aalice_dev·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

456
Kkite·2 days agoLevel 6

This is not a Ubuntu problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

491
Rran_bo·2 days ago

This is not a Ubuntu problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

124
Lli_ming·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

218
Aalice_dev·just now

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

216
RraseMod·1 hour ago

One counter-example: below Ubuntu 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

205
Oops_wang·2 days agoedited

Saved. I am reworking this area this week — this saves a lot of wrong turns.

283
Lli_ming·2 days agoedited

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

5
Lli_ming·3 minutes ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

250
Sslow_query·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

162
Aalice_devMod·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

457
Hhuang_ke·just nowLevel 6

Saved. I am reworking this area this week — this saves a lot of wrong turns.

55
Ttang_hao·2 days agoedited

One counter-example: below Ubuntu 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

449
Llinlin·2 days agoLevel 6

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

276
Hhuang_keMod·2 days agoLevel 6

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

159
Zzhu_zong·2 days agoLevel 6

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

73
Ttang_hao·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

118
Bbob_chen·2 days agoLevel 6

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

241
Cchen_devOP·yesterdayLevel 6

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

156
Nnikic·12 minutes agoLevel 6

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

65
Nnikic·2 days agoLevel 6

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

44
Oops_wang·2 days ago

I just read the Ubuntu source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

81
Kkite·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

28
Sswoole_lee·2 hours ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

104
Lli_ming·3 minutes ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

17
WwinterOP·12 minutes ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

208
Sslow_query·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

328
Sswoole_lee·2 days agoLevel 6

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

282
Rrase·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

178
Wwinter·2 days agoLevel 6

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

186
LlinlinOP·5 hours ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

23
Oops_wangOP·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

11
Ddev_zhou·yesterday

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

13
Ttang_hao·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

2
Kkite·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

309
Rrase·2 days agoLevel 6

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

114
Aalice_dev·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

54
Wwinter·12 minutes ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

53
Ttang_hao·5 hours agoLevel 6

This is not a Ubuntu problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

319
Rran_bo·2 days agoLevel 6

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

141
Bbob_chen·2 days agoLevel 6

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

85
Bbob_chenMod·just now

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

108
Lli_ming·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

191
Ddev_zhouOP·2 days ago

I just read the Ubuntu source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

122
KkiteOP·2 days ago

I just read the Ubuntu source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

119
Ddev_zhou·1 hour ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

174
Bbob_chenOP·2 days ago

I just read the Ubuntu source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

110
Ddev_zhou·2 days ago

This is not a Ubuntu problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

131
Mmike_xu·2 days agoedited

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

116
Lli_ming·3 minutes ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

109
Mmike_xu·1 hour ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

108
Nnikic·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

97
Bbob_chen·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

92
Mmike_xu·2 days ago

I just read the Ubuntu source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

45
Ddev_zhou·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

89
Sswoole_lee·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

465
Zzhou_yiOP·2 days ago

One counter-example: below Ubuntu 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

26
Mmike_xu·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

145
Aalice_devOP·2 days ago

One counter-example: below Ubuntu 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

1
Rrase·2 days agoedited

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

51
Sslow_query·2 days agoedited

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

32
Aalice_dev·just now

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

33
Ttang_hao·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

23
Zzhu_zongMod·1 hour ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

32
Lli_ming·28 minutes ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

23
Wwinter·2 days ago

This is not a Ubuntu problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

21
Zzhu_zong·28 minutes ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

23
Aalice_dev·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

400
Zzhou_yi·28 minutes ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

19
Cchen_dev·28 minutes agoedited

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

19
Nnikic·28 minutes ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

410
Zzhu_zong·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

329
Aalice_dev·28 minutes ago

One counter-example: below Ubuntu 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

508
Hhuang_ke·2 days agoedited

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

77
Cchen_dev·2 days ago

This is not a Ubuntu problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

462
Kkite·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

34
Zzhou_yi·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

13
Zzhou_yiMod·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

7
Oops_wang·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

7
Nnikic·yesterday

One counter-example: below Ubuntu 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

6
Sslow_query·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

302
Zzhou_yi·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

501
Rran_boOP·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

484
Bbob_chen·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

212
Hhuang_ke·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

6
Rran_bo·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

6
Aalice_dev·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

5
Wwinter·2 days ago

I just read the Ubuntu source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

1
Kkite·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

109

This is the post detail page /en/c/ubuntu/post/p8. Posts and comments are generated deterministically from a seeded PRNG, so the same post always renders the same content and the link can be shared, reloaded and indexed. In production this page reads MySQL for the post, Redis for hot-post caching, and fetches the whole comment tree in a single query on the path column.

See the database schema →