Platform SRE
Role purpose:
- Ensure reliability, operability, and continuous improvement of TD SYNNEX enterprise platforms across hybrid cloud and on‑prem environments.
- Engineering‑driven operations focused on automation, Infrastructure‑as‑Code (IaC), observability, and toil reduction.
- Serve as the L3 escalation for complex incidents; continuously improve platform run posture and readiness for L1/L2 execution.
Core responsibilities:
- Platform reliability (hybrid cloud + on‑prem): Own L3 reliability posture; define SLOs/KPIs; lead operability gates and production readiness; maintain runbooks/SOPs.
- Automation & IaC: Design/build operational automation (health checks, remediation workflows); develop Terraform/Ansible configurations; script with Python (preferred), PowerShell, and/or Bash; integrate with ITSM for auditable self‑service and controlled remediation.
- Incident/problem/RCA (L3): Lead diagnosis, stabilization, and recovery for major incidents; drive problem management, RCA, preventive actions; reduce MTTR/MTTD via better signals, runbooks, and automation.
- Observability standards: Define actionable signals, alert quality, dashboards, logging; tune alerting to reduce noise; run data‑driven operational reviews.
- AIOps enablement: Advance predictive/proactive operations (anomaly detection, trend/capacity analysis); support Python‑based analytics and ML/DL where applicable; industrialize operational intelligence safely.
- Provider enablement (outsourced L1/L2): Equip provider with clear runbooks, training, standard changes, escalation criteria; govern performance and ITSM alignment; drive continuous improvement.
- Collaboration & CI: Partner with Platform Engineering to ensure operable‑by‑design capabilities; feed operational insights into roadmap; mentor peers and promote engineering‑led operations culture.
Required qualifications:
- 5+ years in platform/SRE/operations/platform engineering with production ownership in large‑scale environments.
- Hands‑on hybrid operations (cloud + on‑prem) with strong enterprise cloud fundamentals (compute, networking, storage, identity).
- Production IaC and automation (Terraform, Ansible); scripting with Python/PowerShell/Bash (Python strongly preferred).
- Proven L3 incident troubleshooting and major incident leadership.
- Strong infrastructure fundamentals: networking (including DNS/DHCP concepts), virtualization, storage, Windows Server and/or Linux.
- ITSM experience (incident, problem, change) and ticket‑based operations.
- Azure platform knowledge.
Preferred/valued:
- SRE practices (SLOs, error budgets, postmortems, toil reduction).
- Virtualization and backup/DR operations experience.
- Exposure to containers and DevOps/CI/CD; configuration drift control.
- Python for operational analytics; familiarity with ML/DL for anomaly detection, forecasting, clustering.
- Experience in large, multinational, 24/7 operations.
- Knowledge of AI/agentic approaches and modern automation patterns.
Desired attributes:
- Engineering mindset; automates and standardizes to reduce toil.
- Strong ownership; calm, structured incident leader.
- Clear communicator in global, matrixed environments; effective cross‑team/vendor partner.
- Comfortable across cloud and on‑prem; disciplined documentation; committed to operational excellence.
Key competencies:
- Site reliability/operations engineering
- Automation, IaC (Terraform/Ansible), scripting (Python/PowerShell/Bash)
- Hybrid infrastructure operations
- Incident/problem management, RCA, continuous improvement
- Observability and alert quality management (tool‑agnostic)
- Provider enablement and operational governance
- Data‑driven operations and AIOps‑oriented thinking
At Tech Data, a TD SYNNEX Company, our values guide everything we do: Together, We Own It, We Dare to Go, We Grow and Win, and above all, We Do the Right Thing. These principles shape how we work with each other, our partners, and our communities as