Anneke Visser
Site Reliability Engineer · Chicago, IL · anneke.visser@example.com · linkedin.com/in/annekevisser
Summary
Site reliability engineer with 6 years running high-traffic production systems: SLO programs, incident command, capacity planning, and toil automation. Raised checkout availability from 99.5% to 99.95% and cut MTTR 70% while halving pages per on-call shift. Strong Kubernetes, Prometheus/Grafana, Terraform, and Go/Python automation; incident commander for 40+ SEV1/SEV2s.
Professional Experience
Site Reliability Engineer · Vantage Rail
Apr 2022 – Present
- Raised checkout availability from 99.5% to 99.95% by defining SLOs with product owners, gating releases on error budget, and automating rollback on burn-rate alerts.
- Cut MTTR 70% (48 to 14 minutes) by rebuilding alerting around symptom-based SLO burn rates and attaching a tested runbook to every page.
- Halved pages per on-call shift (11 to 5) by deleting noisy alerts, auto-remediating the top 6 page causes, and driving a monthly alert-review ritual.
- Survived a 9x Black Friday traffic peak with zero customer-facing incidents by load-testing to breakpoint, pre-scaling with HPA/cluster-autoscaler, and running a game-day program.
- Cut toil 30% of team time to under 10% (tracked quarterly) by automating certificate rotation, failover drills, and capacity reports in Go and Python.
Software Engineer, Infrastructure · Brightpath Logistics
Jul 2019 – Mar 2022
- Eliminated a recurring outage class (cascading retries) affecting 200K users by adding circuit breakers and backpressure across 8 services.
- Cut deployment failures 60% by building a canary-analysis step that auto-halted bad releases on golden-signal regression.
- Built the on-call onboarding program (shadow rotations, runbook standards) that took new engineers from hire to primary in 6 weeks instead of 4 months.
Projects
Open-source: burnrate
- Published a Prometheus rule generator for multi-window SLO burn-rate alerts (Google SRE workbook method); adopted by several teams and 500+ GitHub stars.
Technical Skills
- Reliability engineering: SLOs / SLIs / error budgets, incident command, capacity planning & load testing, chaos engineering / game days, runbook automation, postmortem facilitation
- Platform & tooling: Kubernetes, Prometheus & Grafana, Terraform, AWS (EKS, RDS, VPC), Go, Python
Certifications & Education
- B.S. Computer Science, University of Illinois Urbana-Champaign — May 2019
- CKA: Certified Kubernetes Administrator — Aug 2021