Site Reliability Engineer, Swift Tech

Employer not named by the sourceRemote

MobileDevOps/SREFull Stack

Apply on the company’s site

Frontier is not the employer and does not collect applications.

About this role

Python, Linux, Cloud Computing, Docker, Site Reliability Engineering, CI/CD, Incident Response · Job Description Keep client environments running at 99.9%+ uptime and lead incident response across Swift Tech Co.'s SRE function, the person paged first (via PagerDuty) when an SLO burns, and the person who makes sure it burns slower next time, with real DORA metrics tracking whether that's actually happening Responsibilities • Own SLOs and error budgets for managed client environments • Build and improve alerting with Prometheus, Grafana, and PagerDuty so pages are actionable, not noise • Instrument services with OpenTelemetry and use distributed tracing (Tempo or equivalent) to find root cause fast • Lead postmortems for major incidents and track remediation items to completion • Track and report on DORA metrics (deploy frequency, lead time, change failure rate, MTTR) across client environments • Partner with engineering teams on reliability reviews before, not after, a service ships