Neocloud Service Delivery Manager

๐Ÿข Mirantis ยท all Mirantis jobs
๐Ÿ“ Czechia
๐Ÿ“… Posted 2026-08-23 ยท via Himalayas
๐Ÿท Service-Delivery-Manager,Cloud-Operations-Manager,Infrastructure-Operations-Manager,Technical-Operations-Manager,SRE-Manager,Cloud-Delivery-Manager,Cloud-Delivery-Management,Services-Delivery-Management,Services-Delivery-Manager,Service-Delivery-Management,Cloud-Service-Management,IT-Service-Delivery-Management
Apply on original site โ†—

About the Role

Mirantis is expanding into Neocloud Infrastructure-as-a-Service โ€” managing large-scale infrastructure to a high SLA on behalf of customers running demanding compute workloads, with the scope growing over time to cover the full stack up to the platform layer. We're looking for a Neocloud Service Delivery Manager to lead the engineering team delivering this service, own our incident and outage management process, and make sure our customers consistently get the level of service we commit to.

This is an operational leadership role. You'll manage a team of engineers responsible for infrastructure and, over time, the platform layer running on top of it. You'll own the incident lifecycle end-to-end and act as the escalation point when things go wrong โ€” while building the reporting, process, and team rhythm that reduces how often they go wrong in the first place.
Key Responsibilities
Team Management

-
Manage a team of engineers responsible for large-scale infrastructure and platform operations

-
Own rostering and coverage across a 24x7, globally distributed operating model, ensuring continuous operational coverage

-
Run team performance management: 1:1s, goal-setting, skills development, and performance reviews

-
Identify and close skills gaps within the team as scope grows from infrastructure into platform operations

-
Own onboarding of new engineers into live customer environments

Service Delivery & SLA Management

-
Own service delivery performance against contracted SLAs across your customer portfolio

-
Define, track, and report on service KPIs (availability, MTTR, MTTA, ticket aging/backlog, customer satisfaction)

-
Chair regular service review meetings, both internal and customer-facing, backed by clear data

-
Maintain and continuously improve runbooks, escalation paths, and operational documentation

-
Manage SLA risk and escalate commercial impact to leadership where relevant

Incident & Critical Outage Management

-
Own the incident management process end-to-end: detection, triage, escalation, resolution, and post-incident review

-
Act as the primary escalation point for major and critical incidents/outages, coordinating cross-functional resources and communicating status to customers and leadership

-
Understand the technical detail behind incidents well enough to ask the right questions, challenge root-cause analysis, and make good calls under pressure

-
Run blameless post-incident reviews (RCA/PIR) and track corrective actions through to closure

-
Track incident trends over time to drive proactive reliability improvements rather than reactive firefighting

-
Ensure on-call and escalation rotas are staffed, documented, and regularly tested

Customer & Stakeholder Management

-
Act as a senior operational point of contact for key customers, building trust through consistent, transparent delivery

-
Partner with Sales and Solutions Architecture during onboarding of new customers to confirm delivery readiness

-
Represent service delivery performance and improvement plans in customer business reviews

Process & Continuous Improvement

-
Drive continuous improvement across monitoring, observability, and incident-management tooling to reduce manual effort and improve detection speed

-
Maintain structured onboarding/handover documentation for new customer environments

-
Ensure operational readiness reviews are completed before any new customer environment goes live

What Success Looks Like

-
SLA targets consistently met or exceeded across your customer portfolio

-
Incidents detected, escalated, and resolved within target timeframes, with clear customer communication throughout

-
A stable, well-rostered team with a visible skills growth path as scope expands from infrastructure to platform

-
A shrinking rate of repeat incidents, driven by disciplined post-incident follow-through

-
Customers who trust the team because they can see the data, not just hear reassu

โ† All remote jobs