TensorRT · Beginners · Hiring: remote tensorrt-newbie engineer (full-time, long-term)

708
TEr/tensorrt-newbie·posted by nikic·3 hours agoOpen source

Hiring: remote tensorrt-newbie engineer (full-time, long-term)

Most tensorrt-newbie articles stop at "how to use it" and never cover "when not to use it". This is an attempt at the second half.

What genuinely surprised me was the tail. The average looked great while P99 jumped by an order of magnitude past some threshold. The cause was not tensorrt-newbie itself but our upstream connection reuse — the load test traffic was too clean and hid the long-tail requests.

On trade-offs, my view is this: if nobody on the team owns this area long-term, do not introduce a second mechanism. With two coexistence you first have to work out which one is even in play when things break, and that costs far more than the performance you saved.

Worth noting: the official docs do cover this, just in a very inconspicuous spot. I only found it reading the source comments, where the author explains the reasoning — roughly "so that it degrades into predictable behaviour in extreme cases".

We also fixed monitoring along the way: replaced average-based alerts with percentiles and split them per endpoint. False alerts dropped by about seventy percent and the on-call rotation visibly cheered up.

353 comments

353 comments

· first 120 loaded
M
Bbob_chen·2 days ago

I just read the tensorrt-newbie source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

509
Lli_ming·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

497
Zzhou_yi·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

488
Oops_wang·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

422
Sslow_query·3 minutes ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

471
Nnikic·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

148
Cchen_dev·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

507
Wwinter·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

48
Aalice_dev·2 days ago

I just read the tensorrt-newbie source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

69
Rran_bo·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

411
Cchen_devOP·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

2
Mmike_xu·28 minutes ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

356
Sslow_query·just now

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

106
Wwinter·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

367
Ddev_zhouOPMod·3 minutes ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

159
Wwinter·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

484
Nnikic·2 days ago

One counter-example: below tensorrt-newbie 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

178
Oops_wang·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

283
Rrase·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

69
Sswoole_lee·3 minutes ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

323
Oops_wang·1 hour ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

216
Wwinter·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

140
Kkite·2 days agoedited

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

264
RraseOP·2 days ago

I just read the tensorrt-newbie source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

183
Rran_bo·2 days ago

This is not a tensorrt-newbie problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

217
Lli_ming·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

192
Cchen_dev·5 hours ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

6
Ddev_zhou·28 minutes ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

207
KkiteOP·2 days agoedited

Saved. I am reworking this area this week — this saves a lot of wrong turns.

41
Sswoole_lee·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

123
Rrase·2 days ago

One counter-example: below tensorrt-newbie 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

43
Ddev_zhou·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

188
Llinlin·1 hour ago

One counter-example: below tensorrt-newbie 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

162
Rrase·just now

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

175
Llinlin·1 hour ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

81
Ddev_zhou·2 hours ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

121
Oops_wang·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

361
Kkernel_panic·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

1
Aalice_dev·2 days ago

One counter-example: below tensorrt-newbie 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

36
Rrase·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

81
Bbob_chen·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

163
Kkite·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

163
Kkite·2 days ago

One counter-example: below tensorrt-newbie 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

149
Sslow_query·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

149
Llinlin·2 days ago

I just read the tensorrt-newbie source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

147
Cchen_devOP·1 hour ago

This is not a tensorrt-newbie problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

85
Aalice_dev·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

138
Bbob_chen·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

481
Sswoole_lee·2 days agoedited

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

137
Cchen_dev·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

7
Oops_wang·2 days ago

This is not a tensorrt-newbie problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

250
Zzhu_zong·28 minutes ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

365
Zzhou_yi·2 days agoedited

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

42
RraseMod·2 days ago

This is not a tensorrt-newbie problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

7
Bbob_chen·28 minutes ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

129
Zzhu_zong·2 days ago

I just read the tensorrt-newbie source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

128
Aalice_dev·2 days ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

49
Rran_bo·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

266
Sslow_query·2 days agoedited

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

33
Ttang_haoOP·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

66
Ddev_zhou·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

5
Wwinter·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

121
WwinterOP·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

50
KkiteOPMod·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

33
Nnikic·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

156
Hhuang_keOP·28 minutes ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

30
Bbob_chen·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

87
Bbob_chen·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

87
Zzhu_zong·2 days agoedited

One counter-example: below tensorrt-newbie 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

84
Oops_wang·just nowedited

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

80
Kkite·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

167
Lli_ming·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

380
Llinlin·2 days ago

This is not a tensorrt-newbie problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

59
Kkernel_panic·12 minutes ago

I just read the tensorrt-newbie source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

55
Sslow_queryOP·1 hour agoedited

One counter-example: below tensorrt-newbie 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

261
Kkernel_panic·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

440
Nnikic·2 days ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

188
Kkite·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

107
Bbob_chenMod·3 minutes agoedited

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

295
Rran_bo·2 days agoLevel 6

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

250
Mmike_xu·2 days ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

175
Sswoole_lee·28 minutes ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

16
Llinlin·28 minutes ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

35
NnikicOP·2 days ago

There is actually a simpler fix that needs no architecture change: move this check up to the gateway and the problem disappears. The cost is one extra lookup at the gateway.

14
Sslow_query·2 days agoedited

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

33
Sswoole_lee·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

15
Rrase·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

262
Llinlin·3 minutes ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

12
Ddev_zhou·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

84
Wwinter·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

12
Rran_bo·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

10
Hhuang_ke·5 hours ago

We have run this in production for two years without hitting it. That said, we never reached this scale, so our experience is not really evidence here.

9
Hhuang_ke·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

389
Ddev_zhou·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

7
Oops_wang·2 days ago

One counter-example: below tensorrt-newbie 7.4 the semantics of that code are different, so do not copy it verbatim. We got burned in staging and rolled back once.

7
Oops_wang·3 minutes ago

Has anyone run a controlled experiment? I did, reducing it to a single variable, and the difference was 4% — within noise. So I suspect the main cause is something else.

6
Bbob_chen·2 hours ago

Worth learning from this debugging approach. We went straight at the logs and took a much longer route.

414
Cchen_devOP·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

117
Llinlin·2 days ago

I just read the tensorrt-newbie source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

38
Rran_boMod·2 days ago

Saved. I am reworking this area this week — this saves a lot of wrong turns.

160
Rrase·just now

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

6
Ttang_hao·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

92
Rran_boOP·2 days ago

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

461
Sslow_query·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

31
Lli_ming·2 days ago

This is not a tensorrt-newbie problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

126
Hhuang_ke·2 days ago

This is not a tensorrt-newbie problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

297
Oops_wang·2 days ago

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

99
Nnikic·5 hours ago

This is not a tensorrt-newbie problem, it is a usage problem. The docs say this API is not thread-safe and you must lock around it yourself.

5
Bbob_chen·2 days ago

Can you give a minimal reproduction? I ran it locally for ten minutes and could not reproduce on macOS with the latest version.

111
Rrase·2 days ago

A question: what changes in a container with a 512Mi memory limit? That is how we run it in production.

27
Zzhu_zong·2 days agoedited

Thanks for sharing real numbers — far more useful than the articles that only cover concepts.

1
NnikicOP·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

376
Zzhu_zong·2 days ago

This matches what we see in production. We only hit it past 3k QPS; the earlier load tests showed nothing — the test traffic was too clean, with no long-tail requests.

2
Bbob_chen·2 hours agoedited

Saved. I am reworking this area this week — this saves a lot of wrong turns.

2
Sswoole_lee·5 hours ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

1
Aalice_dev·2 days ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

1
Aalice_dev·28 minutes agoedited

Agreeing with the above. One addition: with this option enabled the GC count in your metrics doubles, so adjust the alert threshold at the same time or it will keep firing.

98
Mmike_xu·2 hours ago

I just read the tensorrt-newbie source — the author actually explains the reasoning in a comment, roughly "so that it degrades into predictable behaviour in extreme cases".

96
Oops_wang·28 minutes ago

Sharing our numbers, 8 cores 16GB, same scenario:

| Concurrency | P50 | P99 |
|---|---|---|
| 200 | 12ms | 88ms |
| 500 | 31ms | 340ms |

P99 clearly collapses at 500 concurrency, which lines up with your knee point.

1
Nnikic·2 days ago

I see point 3 differently. The trade-off depends on your read/write ratio: read-heavy with little writing means caching actually widens the inconsistency window.

1

This is the post detail page /en/c/tensorrt-newbie/post/p1. Posts and comments are generated deterministically from a seeded PRNG, so the same post always renders the same content and the link can be shared, reloaded and indexed. In production this page reads MySQL for the post, Redis for hot-post caching, and fetches the whole comment tree in a single query on the path column.

See the database schema →