Lead Site Reliability Engineer
About us
Intellum is the leader in corporate education technology and powers the largest, most successful customer, partner, and employee learning programs in the world. Large brands and fast-moving companies like Google, Meta, Amazon, Walmart, Xero, Atlassian, Mailchimp, Airbnb, Stripe, and TikTok rely on Intellum to engage and educate the audiences they touch.
We have always been a “remote first” company and are proud to have team members located all over the world. We value Curiosity, Creativity, Perseverance, and Kindness and strive to demonstrate these core values every day. Our culture is very important to us. We invest in our people in fun and exciting ways, including personal development budgets and an annual all-company retreat that is focused less on work and more on human connections. We are in growth mode, and our “smart growth” approach ensures that we will continue to scale our company effectively.
The Lead Systems Engineer is a senior individual contributor responsible for the reliability, scalability, and modernization of Intellum 's platform infrastructure. Intellum serves large enterprise customers with demanding availability expectations, and this role will help shape the technical direction for how our platform runs, deploys, and scales.
This is a highly hands-on role with significant ownership across infrastructure architecture, cloud environments, deployment systems, observability, and platform reliability. The Lead Systems Engineer will also provide technical leadership across the Systems Engineering function through architecture guidance, mentorship, knowledge sharing, and strong operational standards.
A key focus of this role is continuing to modernize the platform toward portable, container-orchestrated infrastructure, improving deployment and observability capabilities, and maintaining an architecture that can operate effectively across multiple cloud providers.
Responsibilities
- Own and drive key infrastructure modernization initiatives, including the continued evolution from legacy compute environments toward modern, container-orchestrated infrastructure while maintaining reliable service for enterprise customers.
- Design and maintain infrastructure as code across multiple cloud providers, ensuring infrastructure decisions support portability, maintainability, and long-term scalability.
- Improve the reliability and maturity of Intellum 's CI/CD systems and deployment tooling so releases are efficient, observable, and recoverable.
- Provide technical leadership across the Systems Engineering team through mentorship, architecture guidance, knowledge sharing, and support for strong engineering practices.
- Establish and evolve SLI and SLO practices, along with the monitoring, alerting, and load-testing capabilities needed to support platform reliability.
- Participate in and provide leadership during platform incidents, including troubleshooting, root cause analysis, and follow-through on corrective actions.
- Drive visibility into cloud infrastructure costs and incorporate cost considerations into architecture and infrastructure decisions.
- Improve developer experience by evolving the infrastructure and tooling engineers depend on, including development environments, deployment workflows, and production feedback loops.
- Partner closely with Security and Engineering teams on access controls, infrastructure hardening, compliance requirements, and secure infrastructure practices.
- Identify operational and infrastructure risks early, recommend priorities, and help drive the technical roadmap for the Systems Engineering function.
- Contribute to the continued development of the Systems Engineering team and function, including mentoring engineers and helping build strong technical practices as the organization evolves.
- Perform other duties as assigned.
Required Skills
- 8+ years of hands-on experience in infrastructure, DevOps, platform engineering, site reliability enginee