[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes

🏢 Software Mind · all Software Mind jobs
📍 Canada
📅 Posted 2026-08-16 · via Himalayas
🏷 Site-Reliability-Engineering,DevOps-Engineer,Platform-Engineering,Cloud-Engineer,Kubernetes-Engineering,Senior-Kubernetes-Engineer,Senior-Site-Reliability-Engineer,Senior-Kubernetes-OpenShift-Engineer,Senior-Kubernetes-Platform-Architect,Senior-SRE-Engineer,Kubernetes-Operations-Engineer,Senior-Site-Reliability-Engineering-Architect
Apply on original site ↗

About the Role

This is a Senior SRE role supporting production reliability for a Kubernetes-based UI service / AI experience framework stack.

This is not general infrastructure, and it is not a front-end developer role. The strongest candidates will have production SRE experience across Kubernetes operations, observability, Node.js runtime troubleshooting, JVM / Java service troubleshooting, Splunk, and incident ownership.
What You’ll Do

- Support the deployment, operation, and reliability of production services running on Kubernetes.

- Monitor service health and investigate production incidents across distributed applications.

- Participate in on-call support, incident response, root cause analysis, postmortems, and reliability improvements.

- Troubleshoot application runtime, networking, and service-to-service issues in collaboration with engineering teams.

- Support CI/CD, GitOps-based deployments, observability, and production monitoring.

- Work within a client-directed backlog and established priorities.

Required Qualifications

- 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering , or a closely related role, including strong recent hands-on experience supporting Kubernetes-based production services.

- 3+ years of hands-on production Kubernetes experience strongly preferred. Kubernetes production operations, including deployment, scaling, rollout / rollback, resource tuning, and service-to-service troubleshooting

- Strong production incident response experience, including on-call, runbooks, postmortems, and paging hygiene

-
Splunk experience for log aggregation, search, and production troubleshooting

-
Prometheus and Grafana experience, specifically building alert rules and dashboards, not only using existing dashboards

-
CI/CD and infrastructure-as-code for containerized deployments, including Helm and GitOps tools such as ArgoCD or Flux

- Strong Linux and networking fundamentals, including DNS, load balancing, TCP / HTTP, HTTP/2, and Kubernetes networking

-
Production troubleshooting experience across Node.js and JVM/Java services, with strong depth in at least one runtime environment. Experience may include Node.js heap snapshots, CPU profiling, event-loop and memory analysis, as well as JVM GC log analysis, thread dumps, JVM tuning, and Java service latency investigation.

- Service-to-service authentication experience, including mTLS, certificate rotation, certificate format conversion, and JWT-based service authentication

Nice to Have

- Web Components / Lit experience, to perform first-level debugging of UI-related issues

- Server-side rendering or isomorphic runtime experience

- Canary rollout / multi-version production operations

- Distributed tracing and request-context correlation

- KEDA or event-driven autoscaling

- Experience with enterprise platform integration layers

What We Offer

- Competitive salary and laptop

- Professional development and training opportunities

- Work with cutting-edge cloud and container technologies

- Flexible work arrangements and collaborative team environment

- Impact on organization-wide digital transformation initiatives

We are Software Mind , an awesome team of engineers who are ready to ramp up any top-notch company’s projects! Our aim? To always be one step ahead. Become part of a multicultural company in constant growth with an excellent work environment certified by Great Place To Work!
About the Client

Our client is a leading enterprise software company building highly scalable cloud-native platforms used by organizations around the world. Their engineering teams focus on delivering reliable, secure, and high-performing services while embracing modern DevOps, Kubernetes, and cloud technologies.

You will join a team responsible for ensuring the stability, reliability, and operational excellence of a critical UI service running in production.

Contract Duration: Initial cont

← All remote jobs