Staff Site Reliability Engineer

🏒 Visa · all 12 jobs
πŸ“ Brazil
πŸ“… Posted Sep 18, 2026 Β· via Himalayas
🏷 Site Reliability Engineering, Chaos Engineering, Reliability Engineering, DevOps Engineer, Cloud Engineer, Staff Site Reliability Engineer +8 more
Apply on original site β†—

About Us
Visa is a world leader in payments technology, facilitating transactions between consumers, merchants, financial institutions and government entities across more than 200 countries and territories, dedicated to uplifting everyone, everywhere by being the best way to pay and be paid.

At Visa , you'll have the opportunity to create impact at scale β€” tackling meaningful challenges, growing your skills and seeing your contributions impact lives around the world.

Join Visa and do work that matters – to you, to your community, and to the world. Progress starts with you.
Job Description

The Staff Site Reliability Engineer is responsible for architecting, implementing, and maintaining solutions that ensure Visa ’s application services operate with high availability and reliability. This role contributes by helping define the reliability expectations that services must meet, guiding application teams through resilient architecture and design choices, and creating repeatable evidence that critical services can tolerate realistic failure modes. Because the team is still small, the candidate must be able to work independently, lead complex technical work with limited guidance, and influence application-owning teams without direct authority.

All roles require digital fluency, including the ability to work with emerging technologies such as Generative AI tools (e.g. ChatGPT, Microsoft Copilot) to support everyday work.

The Reliability & Resilience Engineering squad works to improve the reliability and resilience of the services sold to customers. The team provides subject matter expertise in chaos engineering, reliability standards, resilience design, observability, and continuous reliability improvement. This consultant-level role is expected to operate as a high-autonomy technical owner within the function, creating clarity from ambiguity and translating business reliability goals into practical engineering standards, experiments, technical guidance, and improvement plans.

Key Responsibilities:

- Own complex resilience engineering work across priority services.

- Design, plan, conduct, and report on controlled chaos experiments and game days.

- Establish repeatable chaos testing patterns.

- Define reliability and resilience standards.

- Translate incidents, observed failure modes, SLO misses, and experiment findings into actionable engineering improvements.

- Provide architecture and system-design guidance to application teams, especially around circuit breakers, load shedding, timeouts, retries, dependency isolation, graceful degradation, observability, alerting, and failure containment.

- Author technical design documents, experiment plans, standards, post-experiment reports, and improvement proposals.

- Mentor engineers through design reviews, technical guidance, and hands-on support to raise the reliability capability of the teams they work with.

- Deliver documented resilience standards, reusable experiment templates, clear production readiness criteria, completed pilot experiments, prioritised remediation backlogs, and measurable improvements in the resilience posture of selected customer-critical services.

This is a remoteposition. A remote position does not require job dutiesbeperformed within proximity of a Visa office location. Remotepositions maybe requiredto be present at a Visa office with scheduled notice.

Qualifications

Basic Qualifications:

- 5+ years of relevant work experience with a Bachelor’s Degree or at least 2 years of work experience with an Advanced degree (e.g. Masters, MBA, JD, MD) or 0 years of work experience with a PhD, OR 8+ years of relevant work experience.

- Experience in Kubernetes and related technologies such as Helm, ArgoCD, and Terraform.

- Experience in automating complex tasks and processes using programming languages such as Python, Go, and Java.

- Experience in planning and performing chaos experiments using tooling such as Gremlin, Chaos Toolkit, Litmus,

Flights + hotels

This role requires you to be in Brazil. If that means relocating or flying in, it is worth checking fares before you commit to a start date.

Compare flights and hotels β†’

← All remote jobs

Comparing Site Reliability Engineer pay and openings β€” the live median is $128k?All remote Site Reliability Engineer jobs β†’Site Reliability Engineer salary data β†’
Want more like this? Browse every live remote developer role.All remote developer jobs β†’
Get new developer jobs by email
Daily email, only when there's something new. One click to stop.

Get remote developer jobs like this by email

10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.

Similar for you