Sr Staff Site Reliability Engineer

🏢 Palo Alto Networks · all 232 jobs
📍 Bulgaria
📅 Posted Sep 19, 2026 · via Himalayas
🏷 Site Reliability Engineering, DevOps Engineer, Cloud Engineer, Infrastructure Engineering, SRE, Staff Site Reliability Engineer (sre) +5 more
Apply on original site ↗

Our Mission

At Palo Alto Networks ®, we’re united by a shared mission—to protect our digital way of life. We thrive at the intersection of innovation and impact, solving real-world problems with cutting-edge technology and bold thinking. Here, everyone has a voice, and every idea counts. If you’re ready to do the most meaningful work of your career alongside people who are just as passionate as you are, you’re in the right place.
Who We Are

In order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have. If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us!
This role is remote, but distance is no barrier to impact. Our hybrid teams collaborate across geographies to solve big problems, stay close to our customers, and grow together. You will be part of a culture that values trust, accountability, and shared success where your work truly matters. Job Summary

Join a team of senior engineers operating in a large-scale, multi-cloud production environment supporting tens of thousands of enterprise customers worldwide. This is not a typical SRE role — you’ll work at the core of a complex, high-impact system alongside experienced DevOps professionals in a fast-paced, cybersecurity-focused organization.
Your Experience

- Own and operate large-scale, global production environments across multiple cloud providers (GCP, AWS, Azure)

- Actively monitor, investigate, and resolve incidents triggered by automated alerting systems (PagerDuty / Incident Response)

- Drive end-to-end troubleshooting across complex, distributed systems with high context switching

- Design, deploy, and improve monitoring and observability systems (e.g., Prometheus, Grafana) — not just react to alerts

- Collaborate closely with internal teams (CX, CS, Engineering) to ensure system reliability and performance

- Work hands-on with modern DevOps and infrastructure tools including Kubernetes, Terraform, CI/CD pipelines, and GitOps workflows

- Develop and maintain automation and tooling (primarily in Python)

- Gain deep understanding of system architecture and interconnected services

- Contribute to a culture of operational excellence in a high-scale, high-availability environment

- Champion asynchronous communication, documentation, and tooling standards to ensure seamless collaboration across time zones and fully distributed teams.

- On call responsibilities:
Daytime hours (12:00–20:00 CET/CEST, based on candidate location and team coverage needs)

Occasional weekends and holidays (rotation-based)
Qualifications

- 5+ years of experience in SRE roles in production environments at scale

- Strong hands-on experience with Kubernetes and Terraform

- Strong hands-on experience with at least one major cloud platform (GCP or AWS required)

- Experience building and configuring monitoring systems (e.g., Prometheus, Grafana)

- Familiarity with CI/CD and GitOps tools (GitLab CI, GitHub Actions, Jenkins, Flux)

- Proficiency in Python for scripting and automation

- Proven success in a fully remote or distributed team environment, demonstrating strong self-management and time organization.

- Strong troubleshooting and problem-solving skills with a passion for incident handling

- Ability to work in fast-paced environments with high context switching

- Highly responsive, proactive, and ownership-driven

- Strong collaboration and communication skills

- Curious mindset and eagerness to learn

Our Commitment

We’re trailblazers that dream big, take risks, and challenge cybersecurity’s status quo. It’s simple: we can’t accomplish our mission without diverse teams innovating, together.

Flights + hotels

This role requires you to be in Bulgaria. If that means relocating or flying in, it is worth checking fares before you commit to a start date.

Compare flights and hotels →

← All remote jobs

Comparing Site Reliability Engineer pay and openings — the live median is $129k?All remote Site Reliability Engineer jobs →Site Reliability Engineer salary data →
Want more like this? Browse every live remote developer role.All remote developer jobs →
Get new developer jobs by email
Daily email, only when there's something new. One click to stop.

Get remote developer jobs like this by email

10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.

Similar for you