SRE — Certified Kubernetes & Linux SysAdmin
I keep an internal developer platform dependable for 3,000+ engineers at OP — building Backstage plugins, wiring up OpenTelemetry, and running the infrastructure as code. I care about the details that make software boring to operate. This site is deliberately tiny: it ships in a single network round-trip.
- Cut API p95 latency from 800ms to 120ms by reworking the read path.
- Ran a 40-service platform at 99.95% availability across two regions.
- Shipped a zero-downtime data migration for 2B+ rows.