Senior DevOps Engineer, AI Platform
Who We Are
At Network Solutions, we’ve been trusted for decades to help people get online and stay ahead. We’ve been here since the beginning of the internet, and we’re still building for what comes next.
As the original digital identity authority, we help secure domain names, protect brands, and safeguard the infrastructure businesses rely on. We empower our customers to own and manage the assets that define them online, while delivering enterprise-grade security to protect against virtual threats. Our team leverages modern, AI-accelerated tools to streamline how businesses manage their digital presence, making the most of our decades of experience.
The Network Solutions team is here to help online businesses protect what’s theirs and build for tomorrow. That’s why millions trust us to protect their domains, brands, and websites every day.
The impact you’ll make
We are looking for a hands-on Senior DevOps Engineer to build and operate the infrastructure powering our AI platforms, agent runtimes, web applications, backend services, APIs, and shared platform capabilities. Our environment includes a centralized LLM gateway, Python-based agent runtimes, RAG workers, MCP services, asynchronous processing, databases, caches, queues, and observability services.
You will work with AI engineers, application engineers, and architects who define technical designs, then independently translate those designs into reliable, scalable, secure, and observable production infrastructure across Microsoft Azure and Oracle Cloud Infrastructure.
What you’ll do
-
Translate application and platform technical designs into production ready cloud infrastructure with minimal supervision.
-
Design, provision, operate, and troubleshoot Kubernetes environments, primarily Azure Kubernetes Service and Oracle Kubernetes Engine.
-
Support AI workloads including LiteLLM based gateways, Python agent runtimes, RAG workers, MCP services, background workers, and asynchronous processing pipelines.
-
Design and manage ingress and egress networking, load balancers, DNS, TLS, private connectivity, routing, NAT, firewalls, network policies, and service to service communication.
-
Build and operate infrastructure for web applications and backend services, including APIs, databases, caches, queues, scheduled jobs, and event driven workloads.
-
Build and maintain CI/CD pipelines using Jenkins and Bitbucket, integrating Docker, Helm, Kubernetes, ArgoCD, and container registries.
-
Automate infrastructure provisioning and configuration using Terraform, Helm, Kubernetes manifests, Python, Bash, and related tooling.
-
Implement end to end observability using metrics, logs, distributed tracing, dashboards, alerts, health checks, and SLOs.
-
Own production readiness, incident troubleshooting, root cause analysis, scalability, reliability, and infrastructure cost optimization.
-
Create reusable infrastructure patterns that allow engineering teams to launch new services quickly and consistently.
What we’re looking for
-
7 or more years of experience in DevOps, SRE, Platform Engineering, Cloud Infrastructure, or a related role.
-
Strong hands on experience operating production Kubernetes environments and deep knowledge of networking, scheduling, storage, autoscaling, security, and troubleshooting.
-
Strong Microsoft Azure experience, including AKS, networking, identity, storage, and monitoring. OCI experience is preferred, or demonstrated ability to work across cloud providers.
-
Strong cloud networking knowledge across virtual networks, subnets, routing, NAT, load balancers, private networking, DNS, TLS, firewalls, ingress, and egress.
-
Strong experience with Jenkins, Bitbucket, Docker, Terraform, Helm, Kubernetes, and Infrastructure as Code.
-
Proven experience supporting production web applications and backend services, including REST APIs, microservices, background workers, and asynchronous architectures.
-
Hands on experience with da
Get remote developer jobs like this by email
One weekly digest. No spam, unsubscribe anytime.