Spent 20 minutes on a "random" failure in the notification worker. It reproduced 1-in-15 and only in CI. Cause: an N+1 query. Deterministic now. Flaky isn't random — it's a bug you haven't cornered.
A one-line refactor took down the payment service because an unindexed query. Rolled back in 2 min thanks to the kill switch. Every change ships behind a flag now — no exceptions.