Senior DevOps Engineer, Infrastructure & Reliability

๐Ÿข Worth AI ยท all Worth AI jobs
๐Ÿ“ United States
๐Ÿ“… Posted 2026-08-07 ยท via Himalayas
๐Ÿท DevOps-Engineer,Infrastructure-Engineering,Site-Reliability-Engineering,Cloud-Engineer,Platform-Engineering,Senior-Staff-DevOps-Engineer,Senior-Infrastructure-Operations-Engineer,Sr-Staff-DevOps-Engineer,Senior-DevOps-Platform-Engineer,Senior-TechOps-Engineer,Senior-Infrastructure-Software-Engineer,Senior-Infrastructure-Engineer,Infrastructure-DevOps-Engineer
Apply on original site โ†—

Worth AI , a leader in the computer software industry, is looking for a Senior DevOps Engineer to join our Infrastructure team with a singular mission: to make our systems faster, more reliable, and more resilient while making life dramatically easier for engineers shipping software.

This is a hands-on build role. You will spend most of your time writing Terraform, tuning Kubernetes workloads, automating things that are currently manual, and shipping infrastructure changes to production. You'll join a small platform team with an established roadmap and existing patterns, and a strong voice in how the work gets built.

- Implement scalable Infrastructure-as-Code patterns using tools like Terraform to standardize cloud provisioning and reduce configuration drift.

- Own and evolve our Kubernetes platform (EKS or self-managed), ensuring workloads are secure, scalable, and resilient by default.

- Optimize CI/CD pipelines to improve deployment frequency, reduce lead time, and increase confidence in releases.

- Design and enforce secure networking, IAM, and secrets management strategies across environments.

- Improve observability by refining metrics, logs, and tracing using tools like DataDog, ensuring actionable insight into system health.

- Optimize cloud cost efficiency through rightsizing, autoscaling strategies, and architectural improvements.

- Implement disaster recovery planning, backup strategies, and multi-region resilience initiatives.

- Refactor brittle or manually managed infrastructure into automated, testable, and reproducible systems.

- Introduce new infrastructure tooling or architectural shifts and drive adoption through documentation, workshops, and hands-on support.

- Partner with engineering teams to eliminate friction in CI/CD, deployments, and cloud environments.

- Communicate technical trade-offs clearly across engineering and product stakeholders, balancing speed with safety.

Technology Stack
- Cloud & Infrastructure: AWS (EKS, RDS, MSK, S3, Lambda, IAM, VPC)
Containerization & Orchestration: Kubernetes, ArgoCD

Infrastructure-as-Code: Terraform

CI/CD: GitHub Actions

Monitoring & Observability: DataDog
Data & Messaging: PostgreSQL, Kafka, Redis
Languages (as needed): Bash, Python, TypeScript, JavaScript

Requirements

- 8+ years in DevOps, SRE, or infrastructure engineering.

- Proven experience designing and operating production Kubernetes environments at scale.

- Deep hands-on expertise with AWS infrastructure and cloud networking.

- Strong experience building and maintaining Terraform modules across large cloud environments.

- Demonstrated ownership of CI/CD systems and measurable improvement of DORA metrics.

- Experience leading incident response processes and driving meaningful postmortem outcomes.

- Strong understanding of distributed systems, event-driven architectures (Kafka), and database performance (PostgreSQL).

- Proven ability to modernize legacy infrastructure and eliminate manual operational toil.

- Track record of taking a scoped infrastructure project from an ambiguous starting point to production without needing daily direction.

- Demonstrated ability to build trust across teams while raising the reliability bar.

Success Metrics
- System Reliability: Maintain or exceed defined SLO/SLA targets with reduced incident frequency and duration.

- Infrastructure Stability: Reduce production incidents caused by misconfiguration, manual processes, or infrastructure drift.

- Operational Efficiency: Increase the percentage of infrastructure managed through code and automation.

- Cost Optimization: Improve cloud cost efficiency without sacrificing reliability or performance.

Bonus Points (Nice to Have)

- Experience coding applications

- Experience operating high-throughput Kafka clusters (MSK or self-managed).

- Strong background in database performance tuning (PostgreSQL, Redis).

- Experience implementing autoscaling strategies for high-traffic systems.

- Fami

โ† All remote jobs