Senior Site Reliability Engineer (Performance and Scalability)
Your mission is to make DigitalZone able to scale. You will build the platform's capacity to absorb campaign-level traffic spikes, and you will give every engineering team the tools, standards, and practices to load- and failure test their own systems. This is an enablement role at its core: you raise the reliability bar across the org by building capability, not by owning every service yourself.
What you'll do
- Build the platform's scalability foundation: capacity planning, autoscaling, caching, queueing, and graceful degradation designed for large campaign spikes rather than steady-state load.
- Establish load and failure testing as a standard engineering practice, giving teams the frameworks, tooling, and runbooks to test their own services and act on the results.
- Own SLOs, error budgets, and the observability stack (metrics, logs, traces, alerting) across TypeScript, Go, and PHP/Laravel services, and standardize how teams instrument for scale.
- Harden Postgres and AWS infrastructure for performance and availability, and reduce toil through automation and IaC.
- Lead incident response and blameless postmortems, and drive the systemic fixes upstream into design and campaign planning so reliability is built in, not bolted on.
- Partner with engineering teams early on capacity and resilience, acting as the multiplier that makes them self-sufficient at scaling their own systems.
Get remote technology | full-time jobs like this by email
10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.
Similar for you
Get 10 hand-picked remote jobs like this one in your inbox every morning. One email a day, matched to what you browse. No spam, one-click unsubscribe.
No thanks — continue to the application ↗