Xavier L. Neese
Site Reliability Engineer, Streaming media
Dallas, TX · (225) 555 0111 · name@example.com · linkedin.com/in/xavierlneese
Key Qualifications
EXPERIENCE
Seven years in production engineering, four as a site reliability engineer.
INDUSTRIES
Streaming media, live sports technology, cloud hosting.
SPECIALTIES
Service level objectives, error budget policy, incident response, capacity planning.
SYSTEMS
Kubernetes, Prometheus, Grafana, Datadog, PagerDuty, Terraform, GitHub Actions.
CREDENTIALS
Certified Kubernetes Administrator, Cloud Native Computing Foundation.
EDUCATION
Bachelor of Science in computer engineering, University of Texas at Dallas.
Executive Summary
Site reliability engineer with seven years keeping streaming services available through live events, where a bad ten minutes is visible to two million viewers. Owns the service level objectives and error budgets for the playback path and carries the pager for it. Spends the quiet weeks removing the work that made the loud weeks painful.
Signature Achievements
- Defined the playback service level objectives that the product and engineering groups now plan releases against.
- Raised measured availability of the playback path from 99.4 percent to 99.97 percent across two seasons.
- Cut pages waking an engineer overnight from 34 a month to 7 by rewriting alerts to fire on symptoms rather than causes.
- Removed 11 hours of manual work a week by automating the capacity ramp that used to be run by hand before each event.
- Held incident command on a two hour outage affecting 400,000 sessions and produced the review that removed the cause.
Professional Experience
Site Reliability Engineer
Trinity River Media Systems, Dallas, TX 2022 to present
One of five reliability engineers supporting the playback and entitlement services of a streaming platform.
- Owns the service level objectives and error budgets for playback, and reports budget burn to the product group each month.
- Carries a shared on call rotation, acting as incident commander on major events and writing the review afterward.
- Builds monitoring in Prometheus and Grafana so every service level objective has a measurement behind it.
- Runs load and failure exercises before major live events and reports the headroom findings to engineering leadership.
- Tracks toil for the team and spends a fixed share of each sprint automating the tasks that generate it.
Platform Engineer
Bluebonnet Web Services, Irving, TX 2019 to 2022
Supported the hosting platform for 200 customer applications running on Kubernetes.
- Built the alerting and dashboard standard that every customer workload was onboarded to.
- Handled escalations from the support team on production faults across the shared platform.
- Automated cluster upgrades, which took a two day manual procedure down to a scheduled job.
Licensure and Certification
Certified Kubernetes Administrator, Cloud Native Computing Foundation
Education
Bachelor of Science in computer engineering, University of Texas at Dallas, 2019
Core Skills
Service level objectives · Error budgets · Incident command · Kubernetes · Prometheus · Grafana · PagerDuty · Terraform · Capacity planning · Toil reduction