Site Reliability Engineer
Adventure is our Culture. Join a team that celebrates a lifestyle as bold as the terrain we love. At Backcountry, we are rooted in adventure, recognition, and wellbeing on and off the mountain. We spotlight employee stories, celebrate milestones, and offer exclusive outdoor perks. Whether you are at HQ, in a retail store, or remote, you will be part of a team that thrives on energy, exploration, and connection.
Reports to: Gustavo Arguedas (Site Reliability Manager)
Location: Remote - Costa Rica
About the Role
Backcountry's online platform serves as the backbone of our customer experience, and this role exists to ensure its reliability, performance, and scalability. As Site Reliability Engineer , you will partner with software engineering, DevOps, and IT operations teams to optimize systems and applications across a multi-cloud stack.
Within 6โ12 months, you will have contributed meaningful improvements to service resiliency and observability, reduced operational toil through automation, and established yourself as a trusted partner to development and infrastructure teams.
This is a lean team. You will own a lot, move fast, and make decisions with full end-to-end responsibility.
What You'll Do
- Work on service resiliency, performance tuning, and system design across Backcountry's platform
- Drive resolution of critical incidents and ensure fixes are methodically implemented through postmortems
- Leverage AI-assisted engineering tools (Claude Code, GitHub Copilot, MCP-based agents) to investigate, automate, and ship fixes across infrastructure and application repositories
- Reduce toil by designing and implementing automation
- Partner with other Site Reliability Engineers, developers, and architects to evaluate and implement best practices for current and future workloads
- Monitor system health and capacity, taking proactive action to fix problems before they occur
- Collaborate with engineering teams to build, deploy, and support features
- Build and maintain observability (metrics, logs, traces, profiles) and SLI/SLO instrumentation for Backcountry services
- Participate in FinOps initiatives across GCP and AWS, including capacity planning and committed-use discount strategy
- Participate in the on-call support rotation within the SRE team
Required Qualifications
- 3+ years of experience supporting containerized production services, preferably running Kubernetes
- 3+ years of experience with Infrastructure as Code (Terraform, AWS CDK, Ansible, etc.)
- 3+ years of cloud experience operating in Google Cloud Platform and/or AWS (multi-cloud stack; Azure/Entra exposure is a plus)
- Comfortable diagnosing issues and shipping bug fixes directly to application code (not just infrastructure) to keep services reliable and stable
- Comfortable performing deep dives across both infrastructure and application/software git repositories to trace issues end-to-end
- Proficient with AI-assisted coding tools (e.g., Claude Code, GitHub Copilot) and MCP-based agents, used to accelerate investigation, code review, and automation
- Strong knowledge of scripting and programming languages (Bash, Python, and TypeScript/Node.js)
- Experience managing Linux (any major distribution) in production environments
- Excellent understanding of internet application protocols (DHCP, DNS, HTTPS, SSH, etc.)
- Understanding of how DevOps (CI/CD) and SRE practices (SLOs, SLIs) apply to daily work
- Hands-on experience with observability tooling (Grafana, Prometheus, Loki, OpenSearch, or equivalents) and SLI/SLO instrumentation
- Experience with GitOps and Kubernetes packaging (ArgoCD, Helm, Kustomize)
- Proactively track emerging technology trends and developments, evaluating which ones are worth bringing into engineering practice
- Bachelor's degree in computer science or similar, or equivalent experience
- Advanced-level English communication skills, both verbal and written
Preferred Qualifications
- Experience using AI