Senior Site Reliability Engineer (remote within EMEA)

๐Ÿข FYUL ยท all FYUL jobs
๐Ÿ“ Estonia,Poland,Romania,Spain,Ukraine
๐Ÿ“… Posted 2026-08-12 ยท via Himalayas
๐Ÿท Site-Reliability-Engineering,Platform-Engineering,Cloud-Engineer,DevOps,Infrastructure-Engineering,Senior-Site-Reliability-Engineer,Remote-Site-Reliability-Engineer,Staff-Site-Reliability-Engineer-(SRE),Senior-Site-Reliability-Engineering-Architect,Site-Reliability-Engineering-Lead
Apply on original site โ†—

About the team:

Platform Infrastructure builds, operates, and continuously evolves FYUL 's container platform and cloud foundation. We foster a DevOps culture through self-service tooling, enabling product engineering teams to ship reliable, secure, and cost-efficient services as the business scales. The team owns our AWS cloud accounts, Kubernetes platform, cloud networking, observability stack, core databases, CI/CD pipelines, and infrastructure-as-code, and acts as the go-to partner for engineering teams on cloud and DevOps topics.

About the role:

We're hiring a Senior SRE II to join Platform Infrastructure as one of the team's senior individual contributors. At this level, you're the go-to person for our most complex infrastructure problems: you architect and drive large-scale automation and reliability initiatives, set standards other engineers follow, and mentor Associate and mid-level SREs. You'll split your time between hands-on platform work - Kubernetes, AWS, GCP, CI/CD, observability - and technical leadership: proposing designs, reviewing others' work, and helping the team make good build-vs-buy and cost/reliability trade-offs.
Your daily tasks will include:

-
Infrastructure & reliability: Architect and manage highly available, secure, and scalable infrastructure across multiple AWS accounts and environments using infrastructure as code.

-
Design and operate our Amazon EKS clusters, including networking policies, persistent storage, and scaling strategies for containerized workloads.

-
Own and evolve core platform services:cloud networking, Kubernetes, and the databases and messaging systems engineering teams depend on.

-
Automation & infrastructure as code: Drive large-scale automation projects and set standards for using Terraform / Terragrunt and GitOps (ArgoCD) across teams.

-
Lead adoption of automation to reduce manual operational work and keep environments consistent and repeatable.

-
Observability & incident response : Be the go-to person for solving complex, cross-service infrastructure problems.

-
Drive initiatives that improve reliability and observability (Grafana, Prometheus, Loki, Tempo, Mimir) so systems scale with minimal manual intervention.

-
Participate in on-call rotation, lead incident response for production issues, and write clear runbooks, ADRs, and postmortems.

-
Security & cost efficiency : Lead security efforts within the team - IAM, encryption, secure logging - and mentor others on secure infrastructure practices.

-
Audit infrastructure spend regularly and drive cost optimization across the platform (rightsizing, autoscaling, FinOps practices).

-
Collaboration & mentorship: Mentor mid-level SREs, provide detailed feedback, and support onboarding of new team members.

-
Communicate complex technical concepts clearly to both engineers and non-technical stakeholders.

-
Partner with product engineering squads to understand their needs and represent Platform Infrastructure in cross-team initiatives.

Your qualifications:

These reflect the technical bar we hold Senior SRE II's to internally, based on our SRE competency framework and current stack.
- Core technical experience:

-
Solid Linux systems administration background and comfort scripting in Python.

-
Strong AWS knowledge: EKS, IAM (roles, policies, IRSA), VPC networking, RDS, S3, SQS, and familiarity with the Well-Architected Framework; experience in multi-account AWS environments is a strong plus.

-
Hands-on experience operating and troubleshooting Kubernetes (EKS) at production scale, including Helm chart development, CNI networking (we run Cilium), pod networking/IPAM concepts, and container security (ECR, image scanning).

-
Proficiency with Terraform (modules, state management) and ideally Terragrunt for multi-environment management; GitOps experience with ArgoCD.

-
Experience with Postgres, MySQL and/or MongoDB in production scale, including Aurora.

-
CI/CD experience with Jenkins (

โ† All remote jobs