Senior Platform Engineer, Cloud Infrastructure

๐Ÿข Virtasant ยท all Virtasant jobs
๐Ÿ“ United States
๐Ÿ“… Posted 2026-08-21 ยท via Himalayas
๐Ÿท Platform-Engineering,Cloud-Infrastructure,Site-Reliability-Engineering,Kubernetes-Engineering,DevOps-Engineer,Senior-Cloud-Platform-Engineer,Senior-Cloud-Infrastructure-Engineer,Senior-Platform-Engineer,Senior-DevOps-Platform-Engineer,Senior-Platform-Engineering,Senior-Cloud-Systems-Engineer,Senior-Infrastructure-Software-Engineer,Senior-Developer-Platform-Engineer,Senior-Infrastructure-Operations-Engineer
Apply on original site โ†—

Senior Platform Engineer, Cloud Infrastructure

Type: Remote
Coverage: Pacific Hours (8:00 AM โ€“ 5:00 PM PST)
Job Description:

We are looking for a senior engineer to build and operate the cloud-native platform that our products run on. This is a hands-on infrastructure engineering role: you will design and run production Kubernetes platforms, write and maintain the Go services and controllers that extend them, and own the reliability of systems that other engineering teams depend on.

The work spans platform architecture, networking, observability, and production operations. You will write real code: Go services, Kubernetes controllers, custom middleware, but the value you create is measured in platform capability and reliability, not lines shipped. We are looking for someone who is comfortable owning ambiguous, multi-quarter initiatives and driving them to production.
Key Responsibilities:
Platform and Kubernetes Engineering:

-
Design, build, and operate production Kubernetes clusters, including cluster networking, workload isolation, and multi-region topologies.

-
Work with Kubernetes internals: resource quota management, scheduling and cluster behaviour, NetworkPolicy enforcement, and custom controllers or operators.

-
Implement and operate service mesh capabilities: mTLS between services, service-account-level authentication and authorisation, and traffic management across internal and external request paths.

-
Optimise containerised workloads for performance, cost, and resource efficiency.

Software Development in Go, Python or Java:

-
Write, refactor, and maintain production services, controllers, and middleware that extend the platform.

-
Read and contribute to complex existing codebases, including open-source projects that we customise or extend.

-
Build HTTP, REST, and gRPC service interfaces used by internal engineering teams.

-
Write meaningful unit and integration tests, and treat testability as a design property rather than an afterthought.

Reliability and Production Operations:

-
Lead incident response for platform-level issues; investigate root causes and author postmortems that result in durable fixes.

-
Troubleshoot production systems using logs, metrics, traces, and profiling tools.

-
Diagnose and resolve performance and reliability problems across distributed systems.

-
Define and drive SLOs, and build the alerting that makes them actionable.

Infrastructure as Code and Delivery:

-
Own infrastructure as code across the platform, authoring reusable modules and maintaining them as the platform evolves.

-
Build and improve CI/CD and GitOps delivery workflows so that teams can ship safely and frequently.

-
Balance developer velocity against reliability, security, and compliance requirements.

-
Plan and execute cloud migration initiatives, including moving production workloads between cloud providers or environments while maintaining reliability and minimizing downtime.

Observability:

-
Build and maintain metrics, dashboards, alerting policies, and distributed tracing.

-
Instrument services so that failures are diagnosable without a code change.

Collaboration:

-
Partner with product, security, and infrastructure teams to gather requirements and align on architecture.

-
Contribute to design reviews and help set technical direction.

-
Mentor other engineers and raise the standard of engineering practice around you.

Qualifications:
Education and Experience:

-
6+ years of professional experience in software, platform, infrastructure, or site reliability engineering, including significant time operating production distributed systems.

-
Demonstrated experience building and operating production Kubernetes platforms, not only deploying onto them.

-
Production experience writing Go, Python or Java.

-
Experience designing systems from an ambiguous starting point and carrying them to production.

-
Experience migrating cloud services, including planning an

โ† All remote jobs