Principal Site Reliability Engineer

🏢 Tandem Diabetes Care · all 9 jobs
📍 United States
💰 USD 165,000 - 185,000 / annual
📅 Posted Sep 19, 2026 · via Himalayas
🏷 Site Reliability Engineering, SRE, DevOps Engineer, Infrastructure Engineering, Cloud Engineer, Principal Site Reliability Engineer +5 more
Apply on original site ↗

GROW WITH US:
Tandem Diabetes Care creates new possibilities for people living with diabetes, their loved ones, and their healthcare providers through a positively different experience. We’d love for you to team up with us to “innovate every day,” put “people first,” and take the “no-shortcuts” approach that has propelled us to become a leader in the diabetes technology industry.

STAY AWESOME:
Tandem Diabetes Care is proud to manufacture and sell the Tandem Mobi system and t:slim X2 insulin pump with Control-IQ+ technology — an advanced predictive algorithm that automates insulin delivery. But we’re so much more than that. Our company’s human-centered approach to design, development, and support delivers innovative products and services for people who use insulin. Because many of our own team members live with diabetes, or have a loved one impacted by diabetes, the work is personal, and we are committed to the cause. Learn more at tandemdiabetes.com
A DAY IN THE LIFE:
The Principal Site Reliability Engineer (SRE) is responsible for the reliability, availability, and performance of the company's production systems. This role leads day-to-day production support and incident response and progressively replaces reactive firefighting with engineered SRE practice: SLOs, observability, on-call design, runbooks, and automation. It also advances infrastructure automation (Terraform/IaC), CI/CD reliability, disaster recovery, and security and compliance readiness in partnership with cross-functional teams. The role works with a distributed team that includes offshore consulting partners and is expected to raise their capability and independence; success is measured as much by what the team can do without the Principal SRE as by what the Principal SRE delivers personally.
The Principal Site Reliability Engineer (SRE)'s at Tandem are also responsible for:
Production Support & Incident Management

-
Leads day-to-day production support: intake, triage, prioritization, escalation, queue health, and change execution.

-
Establishes consistent support practices across a distributed team, including shift handoffs, ticket quality standards, and clear ownership of open issues.

-
Leads incident management end-to-end: incident command, stakeholder communication, and blameless postmortems with corrective actions tracked to closure.

-
Participates in and coordinates response activities for production security incidents, partnering with Security teams to contain threats, restore services, validate controls, and implement corrective actions.

-
Owns on-call strategy: rotation design, escalation paths, alert tuning, and tooling (e.g., PagerDuty, New Relic), with explicit goals of reducing alert fatigue and building coverage that works across time zones. Participates as a senior escalation tier for high-severity incidents.

-
Builds runbooks that standardize response to common failure modes and enable first-line resolution by engineers who did not build the system.

Reliability Engineering & Observability

-
Defines and owns SLIs and SLOs for critical services, using them to guide monitoring, alerting, and reliability priorities.

-
Drives systemic reduction of MTTD and MTTR through better instrumentation, alerting, diagnostics, and automation.

-
Converts recurring support burden into permanent fixes, automation, or documentation rather than absorbing it as ongoing manual work.

-
Establishes and maintains technology currency and lifecycle management practices for production platforms, ensuring cloud services, Kubernetes clusters, operating systems, runtimes, and infrastructure components remain supported, secure, and aligned with organizational standards. Proactively identifies and mitigates End-of-Life (EOL), End-of-Support (EOS), and technology obsolescence risks.

-
Owns business continuity and disaster recovery readiness for production platforms, including backup and recovery strategies, recovery testing, failover capabilitie

Flights + hotels

This role requires you to be in the United States. If that means relocating or flying in, it is worth checking fares before you commit to a start date.

Compare flights and hotels →

← All remote jobs

Comparing Site Reliability Engineer pay and openings — the live median is $128k?All remote Site Reliability Engineer jobs →Site Reliability Engineer salary data →
Want more like this? Browse every live remote developer role.All remote developer jobs →
Get new developer jobs by email
Daily email, only when there's something new. One click to stop.

Get remote developer jobs like this by email

10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.

Similar for you