ML Infrastructure Engineerโœฆ anywhere

๐Ÿข Clera ยท all 101 jobs
๐Ÿ“ Worldwide
๐Ÿ“… Posted Aug 21, 2026 ยท via Himalayas
๐Ÿท ML Infrastructure Engineer, ML AI Infrastructure Engineer, AI ML Infrastructure Engineer, Machine Learning Infrastructure Engineer, AI Infrastructure Engineer, Deep Learning Infrastructure Engineer +2 more
Apply on original site โ†—

About the Role

We're a seed-stage enterprise AI infrastructure company building the context layer that makes AI agents reliable, accurate, and secure for critical business operations โ€” including highly regulated industries like insurance, banking, asset management, and healthcare. Our platform automatically constructs a governed, real-time domain model across all enterprise data, enabling AI agents to make confident, auditable decisions in production.

As an ML Infrastructure Engineer , you'll own the systems that keep our agents running reliably and fast at scale. This is a hands-on production engineering role โ€” focused on real-world impact, not research. You'll design, build, and scale our inference and model-serving infrastructure as concurrency and customer demands grow.
What You'll Do

-
Own inference and model-serving infrastructure end to end โ€” from initial design through production deployment and ongoing scaling.

-
Build and scale systems that enable AI agents to run reliably and efficiently under high and increasing concurrency.

-
Identify and resolve infrastructure bottlenecks in collaboration with ML and platform engineering teams.

-
Optimize systems for latency, throughput, and reliability across cloud-hosted production environments.

-
Drive observability, monitoring, and debugging practices across our production ML stack.

What We're Looking For
Dealbreakers โ€” all required:

-
5+ years of hands-on experience building and operating machine learning inference systems, model-serving platforms, or ML infrastructure in production environments.

-
Demonstrated experience designing and scaling inference-serving infrastructure using tools such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom systems.

-
Proven ability to optimize production ML systems for latency, throughput, and reliability at scale.

Required skills:

-
Experience with containerization and orchestration (Docker, Kubernetes) for deploying and scaling ML workloads.

-
Background in distributed systems that handle high concurrency and dynamic resource allocation under load.

-
Proficiency with monitoring and observability tooling โ€” e.g., Prometheus, Grafana, ELK stack, distributed tracing.

-
Experience deploying and managing ML systems on cloud platforms (AWS, GCP, or Azure).

-
Proficiency in at least one systems or backend language: Python, Go, Rust, C++, or Java.

Nice to have:

-
Experience with knowledge graphs, semantic search, or graph databases (e.g., Neo4j, Amazon Neptune, or similar).

-
Background in real-time inference or low-latency serving requirements.

-
Familiarity with agentic AI systems, autonomous agents, or multi-step reasoning pipelines.

-
Experience with enterprise data infrastructure, data pipelines, or data integration platforms.

Location & Work Arrangement

This is a full-time, on-site role based in San Mateo, CA . On-site collaboration is an important part of how this small, fast-moving team operates. Visa sponsorship is not available for this position.
Why Join

-
Early-stage opportunity with significant ownership and impact โ€” you'll shape foundational infrastructure decisions.

-
Work on genuinely hard distributed systems problems in a production AI context.

-
Small, experienced team with deep ML and enterprise engineering backgrounds.

-
Well-funded at the seed stage with strong institutional backing and a clear enterprise customer focus.

Originally posted on Himalayas

Travel medical insurance

Working from anywhere means medical cover that follows you, not a policy tied to one country. SafetyWing is built for exactly that: travel medical insurance you can start mid-trip and pay monthly.

Get your global insurance today โ†’

โ† All remote jobs

Comparing DevOps Engineer pay and openings โ€” the live median is $110k?All remote DevOps Engineer jobs โ†’DevOps Engineer salary data โ†’
Want more like this? Browse every live remote data science role.All remote data science jobs โ†’
Get new data science jobs by email
Daily email, only when there's something new. One click to stop.

Get remote data science jobs like this by email

10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.

Similar for you