Senior Site Reliability Engineer (remote within EMEA)
About the team:
Platform Infrastructure builds, operates, and continuously evolves FYUL 's container platform and cloud foundation. We foster a DevOps culture through self-service tooling, enabling product engineering teams to ship reliable, secure, and cost-efficient services as the business scales. The team owns our AWS cloud accounts, Kubernetes platform, cloud networking, observability stack, core databases, CI/CD pipelines, and infrastructure-as-code, and acts as the go-to partner for engineering teams on cloud and DevOps topics.
About the role:
We're hiring a Senior SRE II to join Platform Infrastructure as one of the team's senior individual contributors. At this level, you're the go-to person for our most complex infrastructure problems: you architect and drive large-scale automation and reliability initiatives, set standards other engineers follow, and mentor Associate and mid-level SREs. You'll split your time between hands-on platform work - Kubernetes, AWS, GCP, CI/CD, observability - and technical leadership: proposing designs, reviewing others' work, and helping the team make good build-vs-buy and cost/reliability trade-offs.
Your daily tasks will include:
-
Infrastructure & reliability: Architect and manage highly available, secure, and scalable infrastructure across multiple AWS accounts and environments using infrastructure as code.
-
Design and operate our Amazon EKS clusters, including networking policies, persistent storage, and scaling strategies for containerized workloads.
-
Own and evolve core platform services:cloud networking, Kubernetes, and the databases and messaging systems engineering teams depend on.
-
Automation & infrastructure as code: Drive large-scale automation projects and set standards for using Terraform / Terragrunt and GitOps (ArgoCD) across teams.
-
Lead adoption of automation to reduce manual operational work and keep environments consistent and repeatable.
-
Observability & incident response : Be the go-to person for solving complex, cross-service infrastructure problems.
-
Drive initiatives that improve reliability and observability (Grafana, Prometheus, Loki, Tempo, Mimir) so systems scale with minimal manual intervention.
-
Participate in on-call rotation, lead incident response for production issues, and write clear runbooks, ADRs, and postmortems.
-
Security & cost efficiency : Lead security efforts within the team - IAM, encryption, secure logging - and mentor others on secure infrastructure practices.
-
Audit infrastructure spend regularly and drive cost optimization across the platform (rightsizing, autoscaling, FinOps practices).
-
Collaboration & mentorship: Mentor mid-level SREs, provide detailed feedback, and support onboarding of new team members.
-
Communicate complex technical concepts clearly to both engineers and non-technical stakeholders.
-
Partner with product engineering squads to understand their needs and represent Platform Infrastructure in cross-team initiatives.
Your qualifications:
These reflect the technical bar we hold Senior SRE II's to internally, based on our SRE competency framework and current stack.
- Core technical experience:
-
Solid Linux systems administration background and comfort scripting in Python.
-
Strong AWS knowledge: EKS, IAM (roles, policies, IRSA), VPC networking, RDS, S3, SQS, and familiarity with the Well-Architected Framework; experience in multi-account AWS environments is a strong plus.
-
Hands-on experience operating and troubleshooting Kubernetes (EKS) at production scale, including Helm chart development, CNI networking (we run Cilium), pod networking/IPAM concepts, and container security (ECR, image scanning).
-
Proficiency with Terraform (modules, state management) and ideally Terragrunt for multi-environment management; GitOps experience with ArgoCD.
-
Experience with Postgres, MySQL and/or MongoDB in production scale, including Aurora.
-
CI/CD experience with Jenkins (