Senior Software Engineer, Infrastructure

๐Ÿข Stream ยท company page
๐Ÿ“ Remote โ€” Worldwide
๐Ÿ’ฐ CA$155,000 - CA$200,000 / year
๐Ÿ“… Posted Sep 15, 2026 ยท via RemoteIO
๐Ÿท Go, Kubernetes, Python, AWS, GCP
Apply on original site โ†—
โœ“ Re-crawled against the original posting daily โ€” dead and stale listings expire automatically.

Senior Software Engineer, Infrastructure The role
We are hiring a Senior Software Engineer to help rebuild the platform underneath Stream. Over the next year the infrastructure team is moving from AWS to GCP, moving onto Kubernetes, and relocating 35 to 40 Postgres shards off managed RDS to self-hosted, while the platform keeps serving billions of API requests a month. You will own parts of that outright.
This is a small, senior team without the support structures of a large organisation. You will write code most of the time and make infrastructure calls on your own. Success looks like systems that scale predictably under load, cloud spend that falls per unit of traffic, and migrations that land without incident.
This is a full-time job opening based in Toronto (3 days hybrid).
About Stream
Stream powers real-time Chat , Video , Activity Feeds , and AI Moderation for billions of end-users across thousands of apps, from Strava and Bumble to eBay and Patreon. Our platform processes billions of API requests per month and supports applications with millions of concurrent users, while delivering highly reliable, low-latency services and a great developer experience.
What you will do
- Design, build and operate infrastructure for real-time systems carrying millions of concurrent connections and billions of monthly API requests.

- Drive Kubernetes end to end: cluster architecture, workload design and the migration of existing services. You will be designing clusters, not operating someone else's.

- Re-architect workloads as part of the AWS to GCP migration, for cost and performance rather than a lift and shift.

- Own cloud cost and efficiency work: find the levers, measure them against real spend and utilisation data, and show what moved.

- Write production Go and Python: internal services, platform tooling and automation that change how product and SDK engineers deploy, observe and debug.

- Lead post-migration tuning and capacity planning, closing the loop between the architecture you chose and what production actually does.

- Work with backend, video and moderation engineers on system design, reliability targets and tradeoffs that cross service boundaries.

- Take part in on-call, incident response and root cause analysis, and turn what you find into durable fixes.

What we are looking for
- 5+ years in infrastructure, platform, DevOps or SRE engineering, with clear depth in infrastructure over application development.

- A software engineering background. You have built systems, not only configured them. Production coding experience in Go or Python. Scripting-only backgrounds are not a fit.

- Kubernetes at meaningful production scale, past operations: you have driven cluster strategy, designed workloads, or led a migration, and you have tuned what came out the other side for cost and efficiency.

- Cloud cost or efficiency optimisation you personally led on AWS or GCP, with an outcome you can put a number on. FinOps practice is a plus.

- Direct experience running high-scale, high-load production systems.

- Strong cloud fundamentals across networking, compute, storage and IAM, and the habit of asking why a system behaves the way it does instead of accepting the default.

- Comfortable in a small team: leading a project and reviewing a PR in the same week.

- AI tooling already in your engineering workflow. Applied use, not familiarity.

Bonus points
- Both AWS and GCP, and migration experience between providers.

- PostgreSQL at scale: sharding, replication strategy, partitioning tradeoffs, ideally self-hosted.

- Real-time systems: WebSockets, WebRTC, streaming or other persistent-connection workloads.

- The wider stack: CockroachDB, Redis, Terraform, and a Prometheus-based observability stack.

- An API-first or infrastructure company at scaleup stage.

- Open source contributions to infrastructure or platform tooling.

- Writing or talks on cloud, platform or distributed systems.

- Formal FinOps practice,

โ† All remote jobs

Get new remote jobs like this by email
Daily email, only when there's something new. One click to stop.

Get remote software development jobs like this by email

10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.

Similar for you