SRE & Reliability
Reliability work starts at the first real user, not before
On a payments platform, Max keyed autoscaling off transaction volume and payment throughput instead of CPU utilization. Peak hit 3x normal load, and payments kept clearing. Latency held under 500ms. A different employer, a different peak: he built the Black Friday autoscaling for a fintech payment platform, and it took 10x normal load. He set each threshold from traffic he had already watched. We turn down teams with nothing in production yet. There is nothing to measure, and an SLO guessed in advance is a number we would both have to pretend to believe.
The rest is drills. Back at the payments platform, Max wrote the disaster-recovery runbooks against 30+ documented failure scenarios, then exercised them every month with automated cross-region failover tests, held to a sub-minute recovery time objective. Nobody should be finding out mid-incident whether the runbook works.