Director, Software Engineering (Infrastructure)

๐Ÿข ServiceTitan ยท all ServiceTitan jobs
๐Ÿ“ United States
๐Ÿ’ฐ USD 263,800 - 395,600 / annual
๐Ÿ“… Posted 2026-08-04 ยท via Himalayas
๐Ÿท Director-of-Software-Engineering,Infrastructure-Engineering,Site-Reliability-Engineering,Cloud-Engineer,DevOps-Engineer,Software-Engineering-Director,Infrastructure-Engineering-Director,Cloud-Infrastructure-Engineering-Director,Director-Of-DevOps-Engineering,Director-Of-Platform-Engineering,Director-of-Engineering
Apply on original site โ†—

Ready to be a Titan?
Reporting to the VP of Infrastructure, this role is crucial to the success of ServiceTitan but more importantly to the tens of thousands of trades businesses across the continent we call customers. Keeping ServiceTitan up and humming at 4-9โ€™s availability and high performance is critical to our mission of serving the trades and enabling tens of thousands of businesses across the continent to operate smoothly. ServiceTitan is a mission critical operating system our customers leverage for operating their businesses like generating leads, booking appointments, dispatching technicians, planning inventory, invoicing, accepting payments, accounting, issuing payroll, capacity planning, closing books and then some. This role owns the operating rhythm, availability, release and performance of our software.

The key responsibilities include :

-
Lead, grow and develop a global SRE team of engineers that is able to provide 24x7 coverage.

-
Operate operations center (OC) as well as incident command & response functions for this mission critical software.

-
Achieve and maintain 4-9โ€™s availability across our fleet.

-
Own release management across core application as well as orchestrate a resilient process across functional microservices.

-
Engage in service capacity planning and demand forecasting, software performance analysis as well as system tuning.

-
Partner with development teams to make sure the applications are production-ready, scalable, reliable, and observable from day zero.

-
Measure and optimize system performance, with an eye toward pushing our capabilities forward, getting ahead of customer needs, and innovating to continually improve.

-
Identifies, develops, implements, and maintains practices that ensure the highest levels of uptime, performance, reliability, and security across the production & pre-production environments.

-
Provides thought leadership in issue resolution regarding internal and external technology matters.

-
Demonstrates a wide-ranging knowledge of the businesses across the enterprise and industry expertise.

-
Participation in the research and proposed solutions driving the stability and reliability of our products improving overall company quality and customer satisfaction

-
Establish and maintain relationships with peers and leaders, act as an internal resource for teams and business units

-
Drive operational best practice adoption across critical services, continually looking to lower operational barriers to achieving improved reliability.

-
Partner closely with peer engineering executives to ensure we operate as a single team and represent Service Titan in the technology community as well as interacting with customers assuring them of our continued commitment to their success

The ideal candidate will bring to the table :

-
10 -15 years of software engineering experience with a minimum of 7 years in leadership capacity of a team of 50+ engineers.

-
7+ years of experience supporting infrastructure and services hosted in AWS/GCP or Azure.

-
5+ years of experience delivering, deploying and managing enterprise applications in the cloud.

-
3+ years developing continuous integration/delivery/deployment pipelines and cloud-centric CI/CD tools.

-
3+ years of experience implementing telemetry and observability intelligence and automated remediation.

-
3+ years as a leader implementing scalability, resiliency, performance and security.

-
3+ years establishing and maturing an SRE practice.

-
Comprehensive knowledge of Azure cloud services and monitoring technologies is a definite plus.

-
Experience with building pre-production performance and testing environments and DR/HA constructs in a cloud substrate are definitely required.

-
Experience with Infrastructure as Code (IaC) using tools such as Terraform, Ansible, etc.

-
Extensive knowledge of containerization technology such as Docker, Kubernetes, etc.

-
Strong knowledge of building CI/C

โ† All remote jobs