Six rules for error handling in offline-first-opensource that I settled on
Some background first. Our setup is offline-first-opensource plus three downstream services, seven figures of daily requests, peaking around nine in the evening.
The first thing was to collapse the variables. We were changing config and upgrading the version at the same time, and afterwards nobody could say which change caused what. We rolled back to moving one variable at a time, re-ran three times, and only then did the curve settle. Tedious, but not skippable.
134 votes total