Director, Live Operations

🏢 Reveleer · all 16 jobs
📍 United States
📅 Posted Sep 20, 2026 · via Himalayas
🏷 Director Of Live Operations, Production Operations Director, Site Reliability Engineering Director, Director Of DevOps, Cloud Operations Director, Live Events Director +4 more
Apply on original site ↗

Director, Live Operations

Remote Opportunity

About Reveleer

Reveleer delivers a unified platform spanning risk adjustment, quality improvement, clinical intelligence, and member management for health plans and provider organizations navigating the complexity of value-based care. Trusted by 80+ customer organizations nationwide, the platform integrates data, analytics, and intelligent workflow automation into one governed system designed to support traceable documentation across diagnoses, quality measures, and submissions. With regulatory expertise and transparent, human-in-the-loop AI at its core, Reveleer supports organizations working to advance care quality, strengthen documentation integrity, and sustain the operational readiness needed to navigate audits with confidence.

Position Summary

The Director, Live Operations, leads the technical production-operations function responsible for the availability, reliability, supportability, and continuous operation of CDM-managed customer data services. This leader owns the 24x7 operating model for production data workflows and is the accountable leader for major operational incidents, service restoration, problem management, and continuous reliability improvement.

The role leads an operations-engineering organization spanning cloud production support, data-pipeline operations, observability, infrastructure automation, workflow recovery, secure data movement, access/connectivity, operational tooling, and automation. This is a hands-on technical leadership role: the Director must be able to understand and challenge engineering approaches across AWS, Terraform/infrastructure-as-code, data pipelines, APIs, job orchestration, monitoring, and production automation.

Key Responsibilities

24x7 Production Reliability & Service Ownership

- Own the 24x7 operational health and reliability of CDM-managed production data services and workflows.

- Serve as accountable leader and escalation owner for Sev-1 and Sev-2 production incidents.

- Define SLAs/SLOs, escalation paths, on-call/coverage models, service-health measures, and operational performance expectations.

- Drive service restoration, communication coordination, and corrective-action follow-through.

Incident & Problem Management

- Own CDM incident and problem-management disciplines, including severity definitions, incident command, escalation, RCA, post-incident review, and corrective actions.

- Track recurring failures and use MTTA, MTTR, availability, incident volume, and recurrence metrics to drive systemic improvement.

Cloud & Infrastructure Operations

- Provide technical leadership for production services operating in AWS and related enterprise environments.

- Partner with Engineering on infrastructure-as-code using Terraform or comparable tooling, including repeatable configuration, deployment, and recovery.

- Guide operational issues involving IAM, networking/connectivity, secure file transfer, storage, compute, logging, monitoring, and cloud dependencies.

DevOps, Automation & Operations Engineering

- Lead an automation-first strategy to eliminate repetitive manual work, fragile handoffs, and key-person dependencies.

- Drive scripting, orchestration, automated validation, job recovery, exception handling, and self-healing patterns where appropriate.

- Partner with Data Engineering on CI/CD, APIs, ETL/data pipelines, file movement, deployment/support patterns, and production automation.

- Apply AI-assisted monitoring, troubleshooting, documentation, and workflow automation where appropriate.

Observability & Operational Readiness

- Establish monitoring, logging, alerting, and operational dashboards that provide actionable visibility into production health.

- Define production-readiness gates for workflows transitioning from implementation or engineering into Live Operations.

- Require current runbooks, SOPs, recovery procedures, escalation paths, dependency maps, ownership, and cross-trained co

Flights + hotels

This role requires you to be in the United States. If that means relocating or flying in, it is worth checking fares before you commit to a start date.

Compare flights and hotels →

← All remote jobs

Want more like this? Browse every live remote operations role.All remote operations jobs →
Get new operations jobs by email
Daily email, only when there's something new. One click to stop.

Get remote operations jobs like this by email

10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.

Similar for you