Spent most of the afternoon on a "random" failure in the notification worker. It reproduced 1-in-20 and only in CI. Cause: an N+1 query. Deterministic now. Flaky isn't random — it's a bug you haven't cornered. #buildinpublic
Cut CI runtime on the notification worker by ~60% with lazy-loading the module. Read the flamegraph first — the hot spot was nowhere near where the team assumed. Measure, then cut.
Prod incident: p99 latency on the checkout flow spiked 40x at 03:00. Root cause: an off-by-one in the cursor. Fix was reordering two calls. Postmortem: add the metric BEFORE the incident. #perf
Prod incident: p99 latency on the sync engine blew past every alert threshold at 03:00. Root cause: an N+1 query. Fix was a one-line fix. Postmortem: idempotency is not optional.
Caught a nasty one in review: the ingest pipeline checked auth but not ownership — classic IDOR, any user could read any record by id. One WHERE clause between "fine" and "breach". Always scope by owner.