Senior Site Reliability Engineer (Cloud and Networking) - Remote

๐Ÿข Akamai Technologies ยท all Akamai Technologies jobs
๐Ÿ“ Poland
๐Ÿ“… Posted 2026-07-30 ยท via Himalayas
๐Ÿท Site-Reliability-Engineering,Cloud-Networking,Platform-Engineering,Infrastructure-Engineering,SRE,Senior-Site-Reliability-Engineer,Staff-Site-Reliability-Engineer-(SRE),Senior-Site-Reliability-Engineering-Architect,Senior-Cloud-Infrastructure-Engineer
Apply on original site โ†—

Do you want to own the reliability of cloud load balancing infrastructure that serves thousands of customers at global scale?

Are you a senior technical leader who can drive solutions across distributed teams while mentoring the engineers around you?
Join our Cloud Networking SRE Team
The Cloud Networking SRE team (CNETSRE) is part of Akamai's Infrastructure Engineering & Operations (IE&O) organization. We design, deploy, and manage the reliability of Akamai's core cloud networking products โ€” including NodeBalancer, our production L4/L7 load balancer, and NLB (Network Load Balancer), our next-generation high-throughput L4 load balancing platform. These products are foundational to the Akamai Cloud Compute platform, serving customer workloads across dozens of global regions.
Partner with the best
As a Senior Site Reliability Engineer on the NodeBalancer and NLB stack, you'll own the operational reliability of two generations of load balancing infrastructure: the production NodeBalancer fleet and the next-generation NLB platform with its distributed forwarding architecture. You'll build and maintain observability frameworks, lead incident response for complex multi-service failures, drive safe deployment practices across phased global rollouts, and mentor SRE II engineers on the team. The load balancing platform is actively evolving โ€” future iterations are expected to move toward container-orchestrated deployments, so your Kubernetes expertise will be directly relevant as this stack grows. You'll work closely with NodeBalancer Engineering, the Product Delivery Team, and peer SRE functions to shape how the NB/NLB stack evolves operationally.

As a Senior Site Reliability Engineer, you will be responsible for:

- Owning the SRE lifecycle for NodeBalancer and Network Load Balancer โ€” from design reviews and pre-rollout readiness assessments through production sign-off and ongoing reliability management

- Designing and implementing SLO/SLI frameworks that reflect true customer experience for L4 and L7 load balancing services, and driving action when error budgets are at risk

- Building and maintaining observability pipelines for NB/NLB infrastructure, including Prometheus metrics from load balancing components and system-level sources, and Grafana dashboards that enable rapid incident triage

- Leading technical incident response for complex NB/NLB failures โ€” BGP/VIP issues, failover failures, data plane degradations, and configuration problems โ€” acting as the technical commander and driving root cause analysis and preventive follow-through

- Developing and automating safe deployment workflows for phased NB/NLB releases, including bake period monitoring, feature flag management, and GO/NO-GO validation across global datacenter rollouts

- Reviewing design documents, product requirement Documents and producing actionable SRE input on operational risks, capacity implications, Day-2 concerns, and product strategy gaps

- Building automation and tooling using Python or Go that reduces operational toil and improves team-wide operational capability

- Mentoring SRE II engineers on the NB team, providing hands-on technical guidance, code/config reviews, and raising the bar for the team's SRE practice

- Participating in an on-call rotation for NB/NLB production systems, responding to incidents and driving resolution for customer-facing load balancing infrastructure

- Participate in a scheduled, daytime-only on-call rotation to spearhead technical incident response and resolve complex NB/NLB failures..

Do what you love To be successful in this role you will:

- Have extensive experience in SRE, platform engineering, or infrastructure engineering, working with large-scale distributed systems

- Demonstrate deep expertise with Linux networking fundamentals โ€” routing, BGP, nftables/iptables, ARP, VXLAN โ€” and comfort diagnosing at the packet level using tcpdump, netstat, and similar tools

- Have hands-on experience with L4/L7 load balanci

โ† All remote jobs