Sr. Infrastructure Operations Engineer

๐Ÿข KLDiscovery ยท all 20 jobs
๐Ÿ“ India
๐Ÿ“… Posted Sep 13, 2026 ยท via Himalayas
๐Ÿท It Infrastructure Operations Engineer, Senior Infrastructure Engineer, Cloud Infrastructure Engineer, Systems Engineer, Infrastructure Architect, Senior Infrastructure Operations Engineer +6 more
Apply on original site โ†—

The Senior Infrastructure Operations Engineer owns the design, administration, and continuous improvement of KLDiscovery 's compute, storage, and cloud infrastructure globally. This role takes end-to-end technical ownership of the physical server, virtualization, block/file/object storage, enterprise backup, and Azure IaaS environments underpinning KLDiscovery 's production client systems, and serves as the primary technical escalation point and design authority within the Compute & Storage team. The Senior Infrastructure Operations Engineer leads infrastructure design decisions, drives standards development, partners with Enterprise Architecture and IT Security on architecture and compliance, and provides mentoring and technical direction to Infrastructure Operations Engineers. This role participates in 24x7x365 on-call rotation.
Key Responsibilities

Compute, Virtualization & Storage Architecture:

-
Hold end-to-end technical ownership of KLDiscovery โ€™s physical server and virtualized compute environments; define and enforce configuration standards, capacity thresholds, and operational procedures; lead the design of significant compute changes, cluster expansions, and platform upgrades; review and approve significant configuration changes before implementation

-
Own the design standards, capacity planning, performance management, and operational procedures for block, file, and object storage environments globally; monitor performance and capacity; identify constraints and drive remediation before they impact production; engage vendor support directly for complex issues

Backup, Recovery & Azure IaaS:

-
Own the enterprise backup platform and recovery strategy โ€” job standards, recovery validation procedures, and RPO/RTO alignment; ensure recovery procedures are documented, tested on a defined schedule, and executable by any team member; manage the formal backup coverage request process for consuming teams

-
Own the administration and governance of Azure IaaS infrastructure; define Azure infrastructure standards in conjunction with Enterprise Architecture; monitor and govern cloud costs; surface anomalies and optimization opportunities proactively; partner with Enterprise Architecture on hybrid cloud architecture direction and roadmap input

OS Standards, Patching & Provisioning:

-
Own Windows Server and Linux configuration standards, patching cadence, and hardening baselines across all infrastructure globally; own the server patching function for all server infrastructure โ€” end-user endpoints are excluded and owned by Enterprise Platforms

-
Define and own the server build runbook, provisioning standards, sizing guidelines, and handoff checklists in conjunction with Enterprise Architecture; ensure the provisioning process delivers a configured, network-connected OS at the domain-join boundary with sufficient fidelity to eliminate rework at handoff

Security, Observability & Escalation:

-
Embed security controls into infrastructure design from inception โ€” network segmentation, least-privilege access, encryption at rest and in transit, and audit logging โ€” in conjunction with Enterprise Architecture and IT Security; support audit and compliance reviews as needed; maintain working familiarity with applicable security and compliance frameworks (ISO 27001, CIS Controls) as applied to infrastructure configuration and access management

-
Partner with the Automation & Observability team to ensure all owned infrastructure is covered by monitoring and alerting; serve as the primary escalation point within the Compute & Storage team for complex or time-critical infrastructure incidents; participate in the 24x7x365 on-call rotation; conduct root cause analysis for significant incidents, lead post-incident reviews and blameless retrospectives, and drive systemic remediation

Continual Service Improvement & Automation:

-
Evaluate existing infrastructure for improvement, consolidation, and modernization opportuni

Flights + hotels

This role requires you to be in India. If that means relocating or flying in, it is worth checking fares before you commit to a start date.

Compare flights and hotels โ†’

โ† All remote jobs

Want more like this? Browse every live remote operations role.All remote operations jobs โ†’
Get new operations jobs by email
Daily email, only when there's something new. One click to stop.

Get remote operations jobs like this by email

10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.

Similar for you