Sr. Site Reliability Engineer - SRE
We are expanding our Site Reliability Engineering (SRE) team and seeking a highly skilled and passionate Senior SRE to join us. As a member of our growing SRE function, you will play a critical role in ensuring the reliability, scalability, and performance of our mission-critical services. This is an opportunity to shape our SRE practices, drive automation, reduce operational toil, and significantly impact our product's operational excellence.
What You'll Do
- Design, implement, and maintain highly available, scalable, and resilient systems that deliver exceptional customer experiences.
- Serve as a subject matter expert for observability, including monitoring, alerting, logging, tracing, dashboards, and synthetic testing.
- Develop robust, maintainable software and self-service tooling to automate operational tasks and improve reliability.
- Identify and eliminate operational toil through automation, process improvements, and systematic problem solving.
- Lead incident response, participate in on-call rotations, and drive blameless post-mortems.
- Define, implement, and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
- Leverage infrastructure as code, GitOps practices, and CI/CD automation using Terraform, Flux, and GitHub Actions.
- Provide reliability expertise during system design reviews and influence architectural decisions.
- Document processes, build runbooks, and mentor engineers across the organisation.
- Leverage AI responsibly to accelerate investigations, improve documentation, reduce toil, and build intelligent operational workflows while maintaining appropriate human oversight, security, and governance.
What You'll Bring
Core SRE Capabilities
- Demonstrated experience operating and improving production systems at scale in an SRE, Production
- Engineering, or Platform Engineering role.
- Ability to rapidly build accurate mental models of complex distributed systems across infrastructure, applications, networking, identity, and observability domains.
- Strong troubleshooting skills with a methodical, evidence-driven approach to incident response and root cause analysis.
- Experience defining and using SLIs, SLOs, and error budgets to guide reliability decisions.
- Excellent written and verbal communication skills.
Technical Domains
Experience across several of the following areas:
- Kubernetes platforms, including Amazon EKS, and service mesh technologies such as Istio.
- Cloud infrastructure and services within AWS.
- Identity and access management systems, including Auth0 and AWS IAM.\
- Networking fundamentals, including DNS, load balancing, routing, TLS, and connectivity troubleshooting.
- GitOps workflows and infrastructure automation using tools such as Flux and Terraform.
- Observability platforms and practices, including metrics, logs, traces, alerting, dashboards, and synthetic monitoring.
- CI/CD systems and engineering workflows.
- Application logging and distributed system debugging.
- Engineering Mindset
A strong SRE:
- Prioritizes service stability and customer impact during incidents.
- Slows down under pressure, gathers facts, and communicates clearly.
- Reduces operational complexity through automation and simplification.
- Identifies and eliminates toil through self-service tooling and process improvement.
- Demonstrates strong scripting and automation instincts.
- Brings a systems-thinking approach to problem-solving.
- Balances short-term remediation with long-term reliability improvements.
Software Engineering for Reliability
- Demonstrated ability to build and maintain automation, tooling, and self-service capabilities using one or more programming or scripting languages such as Python, Go, or Bash.
- Focuses on applying software engineering practices to improve reliability, reduce toil, and enhance developer productivity. Behavioral Expectations
- Calm and effective during high-seve