Sr. Site Reliability Engineer - SRE

๐Ÿข QAD, Inc. ยท all QAD, Inc. jobs
๐Ÿ“ Spain
๐Ÿ’ฐ EUR 70,000 - 115,000 / annual
๐Ÿ“… Posted 2026-08-08 ยท via Himalayas
๐Ÿท Site-Reliability-Engineering,SRE-Engineer,DevOps-Engineer,Cloud-Engineer,Platform-Engineering,Senior-Site-Reliability-Engineer,Staff-Site-Reliability-Engineer-(SRE),Senior-SRE-Engineer,Senior-Site-Reliability-Engineering-Architect,Principal-Site-Reliability-Engineer,Site-Reliability-Engineering-Lead
Apply on original site โ†—

We are expanding our Site Reliability Engineering (SRE) team and seeking a highly skilled and passionate Senior SRE to join us. As a member of our growing SRE function, you will play a critical role in ensuring the reliability, scalability, and performance of our mission-critical services. This is an opportunity to shape our SRE practices, drive automation, reduce operational toil, and significantly impact our product's operational excellence.
What You'll Do

- Design, implement, and maintain highly available, scalable, and resilient systems that deliver exceptional customer experiences.

- Serve as a subject matter expert for observability, including monitoring, alerting, logging, tracing, dashboards, and synthetic testing.

- Develop robust, maintainable software and self-service tooling to automate operational tasks and improve reliability.

- Identify and eliminate operational toil through automation, process improvements, and systematic problem solving.

- Lead incident response, participate in on-call rotations, and drive blameless post-mortems.

- Define, implement, and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.

- Leverage infrastructure as code, GitOps practices, and CI/CD automation using Terraform, Flux, and GitHub Actions.

- Provide reliability expertise during system design reviews and influence architectural decisions.

- Document processes, build runbooks, and mentor engineers across the organisation.

- Leverage AI responsibly to accelerate investigations, improve documentation, reduce toil, and build intelligent operational workflows while maintaining appropriate human oversight, security, and governance.

What You'll Bring
Core SRE Capabilities

- Demonstrated experience operating and improving production systems at scale in an SRE, Production

- Engineering, or Platform Engineering role.

- Ability to rapidly build accurate mental models of complex distributed systems across infrastructure, applications, networking, identity, and observability domains.

- Strong troubleshooting skills with a methodical, evidence-driven approach to incident response and root cause analysis.

- Experience defining and using SLIs, SLOs, and error budgets to guide reliability decisions.

- Excellent written and verbal communication skills.

Technical Domains
Experience across several of the following areas:

- Kubernetes platforms, including Amazon EKS, and service mesh technologies such as Istio.

- Cloud infrastructure and services within AWS.

- Identity and access management systems, including Auth0 and AWS IAM.\

- Networking fundamentals, including DNS, load balancing, routing, TLS, and connectivity troubleshooting.

- GitOps workflows and infrastructure automation using tools such as Flux and Terraform.

- Observability platforms and practices, including metrics, logs, traces, alerting, dashboards, and synthetic monitoring.

- CI/CD systems and engineering workflows.

- Application logging and distributed system debugging.

- Engineering Mindset

A strong SRE:

- Prioritizes service stability and customer impact during incidents.

- Slows down under pressure, gathers facts, and communicates clearly.

- Reduces operational complexity through automation and simplification.

- Identifies and eliminates toil through self-service tooling and process improvement.

- Demonstrates strong scripting and automation instincts.

- Brings a systems-thinking approach to problem-solving.

- Balances short-term remediation with long-term reliability improvements.

Software Engineering for Reliability

- Demonstrated ability to build and maintain automation, tooling, and self-service capabilities using one or more programming or scripting languages such as Python, Go, or Bash.

- Focuses on applying software engineering practices to improve reliability, reduce toil, and enhance developer productivity. Behavioral Expectations

- Calm and effective during high-seve

โ† All remote jobs