Senior Site Reliability Engineer

🏢 Radiology Partners · all Radiology Partners jobs
📍 United States
💰 USD 125,000 - 145,000 / annual
📅 Posted 2026-07-30 · via Himalayas
🏷 Site-Reliability-Engineering,DevOps,Platform-Engineering,Cloud-Engineer,Senior-Site-Reliability-Engineer,Senior-Site-Reliability-Engineering-Architect,Principal-Site-Reliability-Engineer,Senior-Reliability-Engineer,Site-Reliability-Engineering-Lead,Site-Reliability-Engineering-Manager
Apply on original site ↗

General information

Press space or enter keys to toggle section visibility

Job Title

Senior Site Reliability Engineer

Functional Area

Teammate - Information Technology

City

Remote

Work Location Type

Remote

State

Remote

Employment Type

Full-time (30+ hrs/week)/FULLTIME

Description & Requirements

Press space or enter keys to toggle section visibility

Position Description & Requirements

WHO WE ARE AND WHAT WE DO Mosaic Clinical Technologies™ is pioneering a new imaging paradigm—shifting radiology from reactive diagnostics to proactive, decision-driven care. Through our AI-native platform, MosaicOS™, we unite fragmented technologies into a single, intelligent ecosystem designed to help clinical teams work faster, think smarter and deliver better care. Built cloud-native and designed for continuous iteration, MosaicOS™ is engineered to solve the industry’s most pressing challenges. Our platform expands clinical capacity, improves diagnostic quality by combining human expertise with AI, and turns operational complexity into strategic advantage through seamless integration or replacement of legacy imaging systems.

Every AI tool within MosaicOS™ is validated across tens of millions of exams to ensure safe, reliable and meaningful clinical impact. With a ransomware-resistant design and built-in redundancy, our platform is designed to ensure continuity of care and protect both patients and operations.

Mosaic Clinical Technologies™ is building technology that evolves with clinicians to advance clinical excellence and improve patient outcomes. If you’re passionate about shaping the future of medical imaging, join us!

POSITION SUMMARY

- Lead Site Reliability Engineering (SRE) initiatives by defining and improving SLIs, SLOs, error budgets, and reliability standards for clinical-grade applications and platform services.

- Own production incident management, including high-severity incident response, root cause analysis, blameless post-mortems, and corrective action planning.

- Design and maintain observability solutions across logs, metrics, alerting, and distributed tracing using modern monitoring platforms.

- Optimize Kubernetes-based containerized environments, including service mesh, ingress, autoscaling, performance tuning, and capacity planning.

- Develop automation, Infrastructure as Code (IaC), runbooks-as-code, and self-healing systems to improve operational efficiency and reduce manual support effort.

- Partner with engineering teams throughout the SDLC to embed reliability best practices into architecture, code reviews, CI/CD pipelines, release management, and deployment processes.

- Drive disaster recovery, chaos engineering, load testing, and resiliency initiatives to improve system availability, scalability, and reduce MTTR/MTTD.

- Leverage expertise in AWS, Azure, or GCP, Linux, networking, distributed systems, Kubernetes, and observability tools, while supporting compliance requirements such as HIPAA and HITRUST in a healthcare environment.

- Ability to participate in an on-call rotation

DESIRED PROFESSIONAL SKILLS AND EXPERIENCE:

- 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Production Engineering supporting high-availability production systems

- Bachelor's degree in Information Technology, Computer Science, Engineering, or a related field preferred, Master’s preferred.

- Deep expertise in Kubernetes, Linux, and cloud platforms such as AWS (EKS, IAM, VPC), Azure, or GCP.

- Strong hands-on experience with Python, Go, or Bash, Infrastructure as Code (Terraform, CloudFormation), and implementing SLIs, SLOs, and reliability engineering best practices in production environments.

- Expertise with observability and monitoring platforms, including Datadog, Prometheus, Grafana, CloudWatch, and distributed tracing, as well as production incident management, troubleshooting, and performance optimization at scale.

- Experience supporti

← All remote jobs