Lead Site Reliability Engineer - Imunify Reliability Platform (remote-only)

๐Ÿข CloudLinux ยท all CloudLinux jobs
๐Ÿ“ Poland
๐Ÿ“… Posted 2026-08-23 ยท via Himalayas
๐Ÿท Site-Reliability-Engineering,SRE,Reliability-Engineering,Observability-Engineering,DevOps,Site-Reliability-Engineering-Lead,Senior-Site-Reliability-Engineer,Site-Reliability-Operations-Engineer,Site-Reliability-Engineer,Staff-Site-Reliability-Engineer-(SRE)
Apply on original site โ†—

The problem you'd own

Imunify360 is a multi-layer Linux server security suite โ€” WAF, IDS/IPS, malware scanning and cleanup, proactive defence, patch management, reputation โ€” running as an agent on hundreds of thousands of customer servers, backed by a cloud estate of scanning, correlation and signature-delivery services on our own bare metal.

Roughly 70 components currently ship without a defined service level indicator. Some are internal services we can scrape. Many are agent-side subsystems running on machines we do not own, reporting through a heartbeat we designed for something else. There is monitoring, and there are dashboards, and there is no coherent answer to the question "is this component doing its job right now, and how would we know if it stopped?".

We know the cost of that gap precisely, because we recently paid it: a security control was silently disabled across a large fraction of the fleet for 61 days. Every dashboard was green. The telemetry reported a ruleset version but not whether the control that consumed it was switched on, so a configuration change was indistinguishable from a broken updater. Three independent safety mechanisms existed and all three were gated behind the same condition that caused the failure.

Your job is to make that class of failure detectable in hours instead of months, across the whole product line, and to build the system that keeps it detectable as the product changes.

This is a greenfield charter inside a brownfield estate. You are not inheriting an SRE team, an SLO framework or a paging culture. You are defining them, with the engineering leads, and then making them stick.

What you'll do
1. Define what "working" means for ~70 components

- Run SLI definition with squad leads and senior engineers. You facilitate and hold the standard; the owning squad signs the SLI.

- Build the taxonomy this product actually needs, which is broader than availability and latency:

-
Service SLIs โ€” availability, latency, error rate for cloud-side services.

-
Fleet SLIs โ€” heartbeat reachability, version and configuration convergence across the installed base.

-
Control-efficacy SLIs โ€” the differentiator. What fraction of protected units have the control effectively enabled and current , not merely installed. Ruleset generation drift, signature age, scan coverage, enforcement-mode distribution.

-
Delivery SLIs โ€” artifact publish success, rule-to-fleet lead time, hotfix time-to-convergence.

-
Pipeline SLIs โ€” ingest lag, verdict latency, queue age, backlog burn.

- Enforce one non-negotiable design rule: an SLI must be measurable from outside the gate of the thing it measures. If the control being off also switches off the signal that would tell you it is off, the SLI is invalid. This is the lesson of the incident above and it is the reason this role exists.

- Attach an SLO, an error budget and an owning squad to each. Tiering is expected โ€” not every component earns a 99.9% target or a pager.

2. Build the collection system

- Design and build the pipeline that gets these indicators off the fleet and into a queryable store: push-based, sampled, privacy-constrained, and with a cardinality budget you set and defend.

- Extend agent-side and service-side instrumentation where the signal does not exist yet, in Python, Go and Rust, working with the owning squads.

- Consolidate the current sprawl of dashboards, ad-hoc queries and reporting paths into a defensible set of instruments, and retire what does not earn its keep.

3. Build alerting and alert management

- Symptom-based, SLO-anchored alerting with multi-window burn-rate semantics. Not threshold soup.

- A three-tier taxonomy โ€” page / ticket / dashboard โ€” with an explicit rule for what is allowed to page a human at 03:00.

- Every alert ships with an owner, a runbook and a documented failure mode, or it does not ship.

- Alert hygiene as a standing practice: quarterly review, deletion counted as a win, actionable-rate

โ† All remote jobs