Staff Software Engineer - SRE & AIOps

🏢 ServiceNow · all 260 jobs
📍 United States
📅 Posted Sep 20, 2026 · via Himalayas
🏷 Site Reliability Engineering, DevOps Engineer, Cloud Engineer, AI Ops, Infrastructure Automation, Staff Site Reliability Engineer (sre) +6 more
Apply on original site ↗

About the role:

ServiceNow is seeking a Staff Software Engineer - SRE & AIOps to contribute to infrastructure automation, operational resilience, and toil elimination across our hybrid cloud and data center operations. Embedded within the Site Reliability & Database Engineering organization, you will implement automation-first systems that reduce manual intervention, accelerate incident remediation, and enable our global engineering teams to operate reliably at scale.

This role combines solid hands-on technical expertise in Kubernetes, cloud platforms, and DevOps practices with growing technical leadership capabilities. You will contribute to SRE tooling design, develop auto-remediation capabilities, and help establish patterns that maintain ServiceNow 's cloud platform reliability while minimizing operational toil across follow-the-sun global teams.
What you get to do in this role:

- Deploy, operate, and troubleshoot production Kubernetes clusters across hybrid and multi-cloud environments, maintaining operational standards and supporting high-velocity application deployments.

- Implement and maintain closed-loop auto-remediation systems that detect, classify, and resolve transient infrastructure failures, leveraging automation frameworks and machine learning insights to reduce MTTR and on-call burden.

- Contribute to the design and evolution of SRE tooling stack, including monitoring platforms, incident management systems, log aggregation, and observability integrations that support global on-call operations.

- Develop and maintain SLO frameworks, alerting policies, and automated runbooks that empower on-call engineers to resolve issues autonomously while managing alert fatigue.

- Build and maintain Infrastructure-as-Code frameworks and GitOps pipelines that enable reproducible infrastructure deployments across hybrid and multi-cloud environments with security and compliance guardrails.

- Support hybrid cloud and data center operations, including on-premises infrastructure, public cloud environments, and workload optimization across multi-region deployments.

- Contribute to adoption of containerization, microservices, and DevOps patterns across engineering teams, establishing CI/CD best practices and network security controls.

- Support on-call rotation operations and incident response processes across different time zones, helping develop runbooks and contributing to post-incident reviews that drive continuous improvement.

- Share knowledge and mentor junior SRE engineers on reliability patterns, incident investigation techniques, and automation best practices.

- Champion a culture of blameless incident analysis, data-driven decision-making, and continuous improvement through knowledge sharing and documentation.

- Identify and systematically automate repetitive operational tasks, from infrastructure provisioning to incident response, improving team efficiency and capacity.

To be successful in this role you have:

- Kubernetes Proficiency: Solid hands-on experience operating production Kubernetes clusters, including deployment models, pod orchestration, resource management, network policies, and troubleshooting runtime issues.

- Incident Remediation Experience: Demonstrated experience designing and implementing automated remediation systems, including alert automation, runbook development, and self-healing mechanisms.

- Cloud Platform Knowledge: Strong hands-on experience with AWS (EKS, EC2, RDS) and/or Azure (AKS, VMs) or GCP (GKE), with understanding of core SRE-related services.

- DevOps & IaC Skills: Solid experience with Infrastructure-as-Code tools (Terraform, CloudFormation) and GitOps practices.

- SRE Tooling Familiarity: Working knowledge of observability platforms, incident management systems, and log aggregation tools.

- Distributed Systems Understanding: Understanding of distributed system challenges, fault tolerance, and resilience patterns.

- On-Call Operations: Experience participating i

Flights + hotels

This role requires you to be in the United States. If that means relocating or flying in, it is worth checking fares before you commit to a start date.

Compare flights and hotels →

← All remote jobs

Comparing Software Engineer pay and openings — the live median is $130k?All remote Software Engineer jobs →Software Engineer salary data →
Want more like this? Browse every live remote developer role.All remote developer jobs →
Get new developer jobs by email
Daily email, only when there's something new. One click to stop.

Get remote developer jobs like this by email

10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.

Similar for you