Lead Infrastructure & Site Reliability Engineering

🏢 Crum & Forster · all 35 jobs
📍 United States
💰 USD 105,400 - 198,100 / annual
📅 Posted Sep 16, 2026 · via Himalayas
🏷 Site Reliability Engineering, Infrastructure Engineering, Platform Engineering, Azure Cloud Engineering, Cloud Security Engineering, Site Reliability Engineering Lead +7 more
Apply on original site ↗

Crum & Forster Company Overview

Travel Insured International (TII), a Crum & Forster company, is hiring a Lead Infrastructure & Site Reliability Engineer.

Travel Insured International is a leading travel insurance provider with more than 30 years in business. As a key component of our Specialty Business Unit, within the Accident & Health division, TII provides travel protection plans to help each individual travel confidently. Travel Insured International is proud to offer products to consumers and to agency partners of all sizes. We're committed to providing dependable coverage, great value, and end-to-end satisfaction for all customers.
Job Description

This is a hands-on engineering role with strong leadership influence, focused on reliability and platform. You will lead through technical credibility and influence—setting standards, shaping direction, and elevating the reliability capability across engineering—while remaining deeply hands-on. As TII's Lead Infrastructure & Site Reliability Engineer, you will own the run-time reliability, observability, cloud security, and platform engineering that keep our all-Azure environment secure, resilient, and always-on—supporting the platform and business.

This is a build-and-enhance role: you will mature our observability across metrics, logs, and traces; establish disciplined incident command and blameless post-incident practices; codify infrastructure with Bicep and Terraform; harden the security posture of our Azure platform; and build the self-service platform capabilities that let engineering teams move fast safely. You will own reliability, security, and platform hands-on while partnering with engineering architecture, engineering leadership, and security stakeholders. With the AVP, Infrastructure, co-create the future-state vision and roadmap for Cloud, Observability, ITSM, and related domains making it a shared, measurable engineering discipline in support of TII's growth target and expansion into new distribution channels.
What you will do:
Infrastructure Strategy

- Own and evolve TII's Azure infrastructure and reliability roadmap, aligned to the Azure Well-Architected Framework.

- With the AVP, Infrastructure to co-create the future-state vision and roadmap for Cloud, Observability, ITSM, and related domains.

- Define standards for compute, network, and platform services that scale with business growth and new channels.

- Drive cloud cost optimization (FinOps)—balancing performance, resilience, and spend.

- Partner with engineering architecture to ensure infrastructure enables design-time resilience and delivery velocity.

Site Reliability Engineering

- Own run-time reliability across availability, performance, scalability, and capacity for TII's platform.

- Mature and expand the SLO practice—defining SLIs, refining the 99.9% (and higher, where warranted) SLOs, and operating error budgets to balance reliability and delivery speed.

- Lead capacity planning and performance engineering to support the platform's growth.

- Drive operational readiness reviews for new services and major releases.

Observability

- Own and mature the observability platform across the three pillars—metrics, logs, and traces—enhancing Grafana/Prometheus and Azure Application Insights.

- Implement distributed tracing across the GraphQL/REST services to accelerate diagnosis and reduce time-to-detect and time-to-resolve.

- Establish meaningful alerting and telemetry that reduce noise and surface real signals.

- Build reliability dashboards that give teams and leadership clear visibility into service health.

Platform Engineering

- Own the internal developer platform and self-service infrastructure capabilities that enable engineering teams to provision and operate safely.

- Define golden paths / paved-road templates that make the reliable, secure, and compliant way the easy way.

- Establish and champion Infrastructure as Code standards using Bicep and Terraform.

- Improve dev

← All remote jobs

Comparing Site Reliability Engineer pay and openings — the live median is $127k?All remote Site Reliability Engineer jobs →Site Reliability Engineer salary data →
Get new remote jobs like this by email
Daily email, only when there's something new. One click to stop.

Get remote jobs like this by email

10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.

Similar for you