Senior AI Platform Engineer

🏒 Permhunt · all 4 jobs
πŸ“ United States
πŸ“… Posted Aug 20, 2026 Β· via Himalayas
🏷 AI Platform Engineering, Mlops, Cloud Platform Engineering, Site Reliability Engineering, DevOps, Senior AI Platform Engineer +4 more
Apply on original site β†—

Our client is a technology consulting company providing operational and engineering services to the high-tech sector. They support platform and infrastructure teams across cloud and distributed environments, helping operate scalable, reliable, and production-ready technology platforms.
The Role

Our client is looking for a Senior AI Platform Engineer to operate, maintain, and continuously improve production AI platforms running on Kubernetes across on-premise, AWS, and GCP environments .

The role combines AI platform engineering, MLOps, Kubernetes, Python, observability, and production operations , working with environments similar to AI on EKS and Kubeflow-based machine learning platforms.

You will also play a senior role in improving engineering practices, mentoring team members, and helping shape the platform roadmap.
Key Responsibilities

- Deploy platform releases and configuration changes using GitOps and DevOps practices.

- Monitor AI platform and service health through logs, metrics, monitoring, and observability tools.

- Improve platform reliability through automation, operational tooling, observability, and self-service capabilities.

- Participate in incident response, root cause analysis, and 24/7 operational rotations.

- Investigate and resolve user, platform, integration, and configuration-related issues.

- Promote strong standards across platform security, reliability, and operational engineering.

- Mentor junior engineers in Python fundamentals and help develop their MLOps capabilities.

- Drive the adoption of MLOps best practices across the engineering team.

- Identify gaps in tooling, technical capabilities, and processes required to support production-grade AI systems.

- Contribute to the technical direction and ongoing development of the AI platform.

Qualifications

-
3+ years of experience supporting production AI, ML, or data platforms using technologies such as Ray, Jupyter, AWS SageMaker, Kubeflow, or similar platforms .

-
5+ years of experience across the AI/ML lifecycle, including development, deployment, DevOps, or MLOps.

-
5+ years of hands-on Python experience supporting AI/ML workflows, applications, or data engineering pipelines.

- Strong practical experience with Kubernetes , including managed platforms such as AWS EKS or Google GKE .

- Good understanding of microservices architectures and service communication patterns .

- Strong troubleshooting skills across application crashes, resource contention, service latency, performance, and scaling issues.

- Experience analysing logs, metrics, monitoring systems, and service-level KPIs within production environments.

Nice to Have

- Exposure to additional AI and data platforms such as Flyte, Hugging Face, Vertex AI, LangChain, Claude Code, or other AI agent platforms .

- Hands-on automation or scripting experience using Bash or Python .

- Relevant Kubernetes or cloud certifications such as CKAD or AWS certifications .

Originally posted on Himalayas

Flights + hotels

This role requires you to be in the United States. If that means relocating or flying in, it is worth checking fares before you commit to a start date.

Compare flights and hotels β†’

← All remote jobs

Comparing DevOps Engineer pay and openings β€” the live median is $110k?All remote DevOps Engineer jobs β†’DevOps Engineer salary data β†’
Get new remote jobs like this by email
Daily email, only when there's something new. One click to stop.

Get remote jobs like this by email

10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.

Similar for you