Site Reliability Engineer (SRE) – Day Shift

🏢 Peraton · all 113 jobs
📍 United States
💰 USD 104,000 - 166,000 / annual
📅 Posted Sep 19, 2026 · via Himalayas
🏷 Site Reliability Engineer, SRE, DevOps Engineer, Cloud Engineer, Platform Engineer, Senior Site Reliability Engineer +3 more
Apply on original site ↗

Responsibilities

Peraton is seeking a Site Reliability Engineer (SRE) to join a team respobsible for the operational reliability of production systems running in AWS Commercial and AWS GovCloud environments. The ideal candidate has working knowledge of deploying and managing system components in Azure and GCP. The primary production workload runs on Red Hat OpenShift Service on AWS (ROSA)

The SRE partners closely with platform engineers, the security team, and application developers to ensure the infrastructure services are reliable, available, and deployed in a way that meets both developer and security requirements.

Work Location: Remote

Shift Schedule: This is a Day Shift position with working hours from 7am – 3pm Eastern Standard Time (EST)
What you will do:

- Operate and maintain production infrastructure services and applications, ensuring availability and reliability, performance, security, and operational health.

- Monitor services and applications using defined SLIs, SLOs, dashboards, alerts, and other observability tools; continuously improve the detection, diagnosis, and resolution of operational issues.

- Partner with application teams to define application observability requirements and implement appropriate metrics, logs, traces, dashboards, and alerts into the organization's observability tooling.

- Manage production incidents and service disruptions, including on-call response, troubleshooting, service restoration, root-cause analysis, and post-incident corrective actions.

- Execute application and infrastructure releases through established deployment pipelines, including promotion through staging and production, validation, rollback, and release-related troubleshooting.

- Manage the operational lifecycle of deployed infrastructure, including upgrades, patching, configuration changes, maintenance, and technology refreshes.

- Assess and improve service resilience through capacity planning, performance testing, failure-mode analysis, disaster recovery, backup, failover, and recovery testing.

- Identify and address reliability risks and operational technical debt by using reliability metrics, incident trends, capacity data, and service health indicators to prioritize improvements.

- Automate operational activities using an everything-as-code approach to improve consistency, repeatability, testing, deployment, recovery, and operational efficiency.

- Collaborate with platform engineering and application teams to identify operational requirements, provide feedback on reusable infrastructure building blocks, and continuously improve the reliability and operability of the environment.

Qualifications
Basic Qualifications:

- Must be a U.S. Citizen with the ability to obtain and maintain the required Public Trust level clearance.

- Bachelor's Degree and 8 years of experience, or a High School diploma/equivalent and 12 years of experience.

- 7+ years hands-on experience in site reliability engineering, DevOps, or production systems engineering.

- Hands-on experience operating in AWS Commercial and AWS GovCloud, including OpenShift (ROSA) or comparable Kubernetes-based platforms

- Strong infrastructure-as-code experience with Terraform and Ansible/Ansible Tower.

- Experience with CI/CD platforms GitLab and Jenkins, including reliability gating and deployment automation.

- Proficient in Linux and Windows Server administration

- Experience with enterprise observability tools such as Dynatrace, Datadog, Splunk and Open Telemetry.

- Demonstrated ownership of an SLI/SLO and alerting program, including error budgets, alert rationalization, and noise reduction.

- Scripting/automation proficiency in Python, Bash, PowerShell, or Go.

- Experience operating in federal or regulated environments (FISMA, FedRAMP, NIST 800-53).

Preferred Qualifications:

- AWS Solutions Architect, AWS DevOps Engineer, or AWS SysOps certification

- Red Hat Certified Specialist in ROSA, Red Hat Certified System

Flights + hotels

This role requires you to be in the United States. If that means relocating or flying in, it is worth checking fares before you commit to a start date.

Compare flights and hotels →

← All remote jobs

Comparing Site Reliability Engineer pay and openings — the live median is $128k?All remote Site Reliability Engineer jobs →Site Reliability Engineer salary data →
Want more like this? Browse every live remote developer role.All remote developer jobs →
Get new developer jobs by email
Daily email, only when there's something new. One click to stop.

Get remote developer jobs like this by email

10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.

Similar for you