iOS · Open source · The edge cases the ios-opensource docs never spell out

781
IOr/ios-opensource·posted by zhou_yi·1 hour agoTooling

The edge cases the ios-opensource docs never spell out

Some background first. Our setup is ios-opensource plus three downstream services, seven figures of daily requests, peaking around nine in the evening.

Worth noting: the official docs do cover this, just in a very inconspicuous spot. I only found it reading the source comments, where the author explains the reasoning — roughly "so that it degrades into predictable behaviour in extreme cases".

Image placeholder · object storage in production
POST /api/uploads → CDN origin pull
389 comments

389 comments

· first 120 loaded
M
Sslow_queryOP·2 days agoedited

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

492
WwinterMod·2 days ago

I just read the ios-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

518
Bbob_chen·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

501
Kkernel_panic·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

375
Llinlin·3 minutes ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

361
Lli_ming·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

302
Zzhou_yi·just nowedited

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

1
Llinlin·2 hours ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

309
Kkernel_panic·2 days ago

I just read the ios-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

220
Llinlin·2 days agoeditedLevel 6

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

189
LlinlinOP·2 days agoLevel 6

One counter-example: below ios-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

55
Hhuang_ke·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

108
Rrase·2 days agoLevel 6

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

20
Hhuang_ke·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

32
Oops_wang·2 days agoeditedLevel 6

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

335
Wwinter·2 days agoLevel 6

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

20
Ttang_haoOP·just now

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

160
Nnikic·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

1
Rrase·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

457
Aalice_dev·5 hours ago

This is not a ios-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

354
Ddev_zhouOP·2 days agoedited

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

380
Ddev_zhou·2 hours ago

One counter-example: below ios-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

314
Llinlin·2 days ago

This is not a ios-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

3
Kkernel_panic·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

309
Wwinter·3 minutes ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

479
Mmike_xu·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

288
Bbob_chen·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

287
Ddev_zhou·3 minutes ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

277
Lli_ming·2 days ago

I just read the ios-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

251
Mmike_xu·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

360
Zzhu_zongOP·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

16
Mmike_xu·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

251
Cchen_dev·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

11
Rrase·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

224
Aalice_dev·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

223
Oops_wang·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

217
Nnikic·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

194
Bbob_chen·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

1
Hhuang_ke·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

159
Kkite·3 minutes ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

152
Lli_mingOP·just now

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

304
Llinlin·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

232
Kkernel_panic·3 minutes ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

142
Lli_mingOP·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

462
Kkernel_panicMod·3 minutes ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

276
Sslow_queryMod·2 days agoedited

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

52
Oops_wang·28 minutes ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

17
Hhuang_ke·2 days agoLevel 6

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

3
Wwinter·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

1
Ttang_haoOP·2 days agoLevel 6

I just read the ios-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

21
Hhuang_ke·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

1
Sslow_query·yesterday

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

23
Kkernel_panicOP·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

382
Lli_ming·2 days agoLevel 6

One counter-example: below ios-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

96
Ttang_hao·2 days agoLevel 6

This is not a ios-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

76
Hhuang_ke·2 days agoLevel 6

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

60
Mmike_xu·12 minutes agoLevel 6

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

2
Aalice_dev·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

17
Bbob_chen·1 hour ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

79
Bbob_chen·2 days ago

One counter-example: below ios-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

491
Nnikic·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

33
Aalice_dev·2 days agoedited

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

138
Aalice_devMod·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

81
Mmike_xu·2 hours ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

50
Sswoole_lee·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

317
Bbob_chen·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

302
Zzhu_zongOP·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

8
Rrase·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

24
Ttang_hao·2 days agoedited

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

509
Aalice_devMod·2 days agoLevel 6

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

44
Ddev_zhou·2 days agoeditedLevel 6

This is not a ios-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

5
Llinlin·2 days agoLevel 6

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

3
Aalice_devMod·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

13
Bbob_chen·yesterdayLevel 6

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

54
Oops_wang·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

112
Llinlin·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

105
Rran_bo·3 minutes agoedited

This is not a ios-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

101
Bbob_chenMod·28 minutes ago

I just read the ios-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

93
KkiteMod·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

187
Oops_wang·2 days ago

This is not a ios-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

23
Nnikic·2 days ago

One counter-example: below ios-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

10
Ttang_hao·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

90
Mmike_xu·28 minutes ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

90
Sslow_query·2 days agoedited

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

13
Cchen_dev·2 days agoedited

One counter-example: below ios-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

88
Sslow_query·1 hour ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

77
Aalice_dev·2 days ago

I just read the ios-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

157
Wwinter·2 days ago

I just read the ios-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

152
Nnikic·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

50
Rran_bo·2 days ago

This is not a ios-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

112
Rran_bo·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

15
Kkernel_panic·2 days agoedited

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

2
Zzhou_yi·1 hour ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

43
Hhuang_keOP·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

345
Zzhu_zong·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

2
Llinlin·2 days ago

One counter-example: below ios-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

33
Nnikic·2 days agoedited

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

28
Kkernel_panic·28 minutes agoedited

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

27
Cchen_dev·3 minutes ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

211
Ddev_zhouOP·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

233
Zzhu_zong·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

264
Bbob_chen·2 days ago

This is not a ios-opensource problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

26
Ttang_hao·yesterday

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

25
Rran_bo·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

24
Wwinter·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

19
Cchen_dev·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

18
KkiteMod·2 hours ago

I just read the ios-opensource source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

16
Aalice_dev·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

2
Aalice_dev·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

16
Aalice_dev·2 days ago

One counter-example: below ios-opensource 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

227
Sswoole_lee·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

13
Zzhou_yi·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

394
Bbob_chen·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

11
Mmike_xuMod·2 days agoedited

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

115
Kkite·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

11
Sswoole_lee·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

20
Ttang_hao·just now

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

10
Cchen_dev·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

5
Hhuang_ke·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

1
Hhuang_ke·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

1

This is the post detail page /en/c/ios-opensource/post/p4. Posts and comments are generated deterministically from a seeded PRNG, so the same post always renders the same content and the link can be shared, reloaded and indexed. In production this page reads MySQL for the post, Redis for hot-post caching, and fetches the whole comment tree in a single query on the path column.

See the database schema →