Six rules for error handling in agent-design that I settled on
Most agent-design articles stop at "how to use it" and never cover "when not to use it". This is an attempt at the second half.
One last trap: in container environments remember to adjust the memory-related parameters in step. Otherwise the host limit and the process expectation disagree, and the symptom is intermittent, unreproducible failure.
Worth noting: the official docs do cover this, just in a very inconspicuous spot. I only found it reading the source comments, where the author explains the reasoning — roughly "so that it degrades into predictable behaviour in extreme cases".
On trade-offs, my view is this: if nobody on the team owns this area long-term, do not introduce a second mechanism. With two coexistence you first have to work out which one is even in play when things break, and that costs far more than the performance you saved.