Principal Site Reliability Engineer (SRE)

🏢 Symmetrio · all 9 jobs
📍 United States
💰 USD 180,000 - 200,000 / annual
📅 Posted Sep 16, 2026 · via Himalayas
🏷 Site Reliability Engineering, Cloud Engineer, DevOps Engineer, Healthcare Technology, AWS Engineering, Principal Site Reliability Engineer +5 more
Apply on original site ↗

Symmetrio is recruiting a Principal Site Reliability Engineer (SRE) for our customer, a rapidly growing healthcare technology organization focused on advanced healthcare technology solutions.

This individual will play a critical role in ensuring the reliability, scalability, security, and performance of a mission-critical SaaS platform supporting healthcare providers across the United States. The ideal candidate will possess a unique blend of cloud infrastructure expertise, application troubleshooting experience, production operations leadership, and customer-facing technical problem-solving skills.

The ideal candidate will be equally comfortable investigating application-level issues, troubleshooting AWS networking and infrastructure, leading production incident response efforts, and collaborating with development teams to improve operational excellence.

This is a fully remote position with a salary range of $180,000–$200,000 , depending on experience.
Responsibilities

- Serve as the primary technical owner for production reliability across U.S. customer environments.

- Investigate and resolve complex issues spanning web applications, APIs, backend services, data pipelines, cloud infrastructure, and customer integrations.

- Lead production incident response efforts, coordinating cross-functional teams to restore service and minimize customer impact.

- Perform root cause analysis and drive corrective actions that improve long-term system stability and resilience.

- Partner with software engineering and platform teams to identify recurring reliability risks and implement sustainable solutions.

- Design, configure, and validate secure customer connectivity solutions including Site-to-Site VPNs, Transit Gateway integrations, routing configurations, and secure network paths.

- Support customer onboarding initiatives by troubleshooting connectivity challenges and ensuring consistent implementation processes.

- Enhance platform observability through improvements in monitoring, logging, alerting, tracing, and operational dashboards.

- Contribute to CI/CD, infrastructure automation, and deployment processes that improve release safety and operational consistency.

- Develop operational tooling that supports incident response, troubleshooting, onboarding, and system monitoring activities.

- Collaborate with engineering leadership to improve cloud architecture, scalability, security, and operational readiness.

- Partner with customer-facing teams to communicate technical issues, remediation plans, and reliability improvements in a clear and effective manner.

- Support compliance, security, and risk management initiatives within highly regulated healthcare environments.

Requirements

- 6+ years of hands-on experience supporting and managing AWS-based production environments.

- 4+ years of experience supporting web applications and backend services (Python/Django experience strongly preferred).

- Experience with AWS networking technologies including VPCs, Site-to-Site VPNs, Transit Gateways, routing, NAT gateways, and security groups.

- Strong experience with Terraform and infrastructure-as-code deployment practices.

- Experience with containerized environments including ECS, Fargate, Kubernetes, or similar technologies.

- Experience building and supporting CI/CD pipelines and release automation processes.

- Familiarity with monitoring and observability platforms such as Datadog, CloudWatch, Sentry, Grafana, or similar tools.

- Experience leading production incidents, outage management, and root cause analysis initiatives.

- Exposure to Windows Server environments, Active Directory, Kerberos, and enterprise infrastructure concepts is preferred.

- Healthcare technology, healthcare SaaS, clinical software, or other regulated industry experience is highly preferred.

- Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related technical field preferred.

Benefits

- Health Care Plan (

Flights + hotels

This role requires you to be in the United States. If that means relocating or flying in, it is worth checking fares before you commit to a start date.

Compare flights and hotels →

← All remote jobs

Comparing Site Reliability Engineer pay and openings — the live median is $128k?All remote Site Reliability Engineer jobs →Site Reliability Engineer salary data →
Want more like this? Browse every live remote developer role.All remote developer jobs →
Get new developer jobs by email
Daily email, only when there's something new. One click to stop.

Get remote developer jobs like this by email

10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.

Similar for you