Senior DevOps Engineer, Infrastructure & Reliability
Worth AI , a leader in the computer software industry, is looking for a Senior DevOps Engineer to join our Infrastructure team with a singular mission: to make our systems faster, more reliable, and more resilient while making life dramatically easier for engineers shipping software.
This is a hands-on build role. You will spend most of your time writing Terraform, tuning Kubernetes workloads, automating things that are currently manual, and shipping infrastructure changes to production. You'll join a small platform team with an established roadmap and existing patterns, and a strong voice in how the work gets built.
- Implement scalable Infrastructure-as-Code patterns using tools like Terraform to standardize cloud provisioning and reduce configuration drift.
- Own and evolve our Kubernetes platform (EKS or self-managed), ensuring workloads are secure, scalable, and resilient by default.
- Optimize CI/CD pipelines to improve deployment frequency, reduce lead time, and increase confidence in releases.
- Design and enforce secure networking, IAM, and secrets management strategies across environments.
- Improve observability by refining metrics, logs, and tracing using tools like DataDog, ensuring actionable insight into system health.
- Optimize cloud cost efficiency through rightsizing, autoscaling strategies, and architectural improvements.
- Implement disaster recovery planning, backup strategies, and multi-region resilience initiatives.
- Refactor brittle or manually managed infrastructure into automated, testable, and reproducible systems.
- Introduce new infrastructure tooling or architectural shifts and drive adoption through documentation, workshops, and hands-on support.
- Partner with engineering teams to eliminate friction in CI/CD, deployments, and cloud environments.
- Communicate technical trade-offs clearly across engineering and product stakeholders, balancing speed with safety.
Technology Stack
- Cloud & Infrastructure: AWS (EKS, RDS, MSK, S3, Lambda, IAM, VPC)
Containerization & Orchestration: Kubernetes, ArgoCD
Infrastructure-as-Code: Terraform
CI/CD: GitHub Actions
Monitoring & Observability: DataDog
Data & Messaging: PostgreSQL, Kafka, Redis
Languages (as needed): Bash, Python, TypeScript, JavaScript
Requirements
- 8+ years in DevOps, SRE, or infrastructure engineering.
- Proven experience designing and operating production Kubernetes environments at scale.
- Deep hands-on expertise with AWS infrastructure and cloud networking.
- Strong experience building and maintaining Terraform modules across large cloud environments.
- Demonstrated ownership of CI/CD systems and measurable improvement of DORA metrics.
- Experience leading incident response processes and driving meaningful postmortem outcomes.
- Strong understanding of distributed systems, event-driven architectures (Kafka), and database performance (PostgreSQL).
- Proven ability to modernize legacy infrastructure and eliminate manual operational toil.
- Track record of taking a scoped infrastructure project from an ambiguous starting point to production without needing daily direction.
- Demonstrated ability to build trust across teams while raising the reliability bar.
Success Metrics
- System Reliability: Maintain or exceed defined SLO/SLA targets with reduced incident frequency and duration.
- Infrastructure Stability: Reduce production incidents caused by misconfiguration, manual processes, or infrastructure drift.
- Operational Efficiency: Increase the percentage of infrastructure managed through code and automation.
- Cost Optimization: Improve cloud cost efficiency without sacrificing reliability or performance.
Bonus Points (Nice to Have)
- Experience coding applications
- Experience operating high-throughput Kafka clusters (MSK or self-managed).
- Strong background in database performance tuning (PostgreSQL, Redis).
- Experience implementing autoscaling strategies for high-traffic systems.
- Fami