ML Infrastructure Engineer

๐Ÿข Clera ยท all Clera jobs
๐Ÿ“ Worldwide
๐Ÿ“… Posted 2026-08-21 ยท via Himalayas
๐Ÿท ML-Infrastructure-Engineer,ML-AI-Infrastructure-Engineer,AI-ML-Infrastructure-Engineer,Machine-Learning-Infrastructure-Engineer,AI-Infrastructure-Engineer,Deep-Learning-Infrastructure-Engineer,ML-Infrastructure-Engineering,AI-ML-Infrastructure-Engineering
Apply on original site โ†—

About the Role

We're a seed-stage enterprise AI infrastructure company building the context layer that makes AI agents reliable, accurate, and secure for critical business operations โ€” including highly regulated industries like insurance, banking, asset management, and healthcare. Our platform automatically constructs a governed, real-time domain model across all enterprise data, enabling AI agents to make confident, auditable decisions in production.

As an ML Infrastructure Engineer , you'll own the systems that keep our agents running reliably and fast at scale. This is a hands-on production engineering role โ€” focused on real-world impact, not research. You'll design, build, and scale our inference and model-serving infrastructure as concurrency and customer demands grow.
What You'll Do

-
Own inference and model-serving infrastructure end to end โ€” from initial design through production deployment and ongoing scaling.

-
Build and scale systems that enable AI agents to run reliably and efficiently under high and increasing concurrency.

-
Identify and resolve infrastructure bottlenecks in collaboration with ML and platform engineering teams.

-
Optimize systems for latency, throughput, and reliability across cloud-hosted production environments.

-
Drive observability, monitoring, and debugging practices across our production ML stack.

What We're Looking For
Dealbreakers โ€” all required:

-
5+ years of hands-on experience building and operating machine learning inference systems, model-serving platforms, or ML infrastructure in production environments.

-
Demonstrated experience designing and scaling inference-serving infrastructure using tools such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom systems.

-
Proven ability to optimize production ML systems for latency, throughput, and reliability at scale.

Required skills:

-
Experience with containerization and orchestration (Docker, Kubernetes) for deploying and scaling ML workloads.

-
Background in distributed systems that handle high concurrency and dynamic resource allocation under load.

-
Proficiency with monitoring and observability tooling โ€” e.g., Prometheus, Grafana, ELK stack, distributed tracing.

-
Experience deploying and managing ML systems on cloud platforms (AWS, GCP, or Azure).

-
Proficiency in at least one systems or backend language: Python, Go, Rust, C++, or Java.

Nice to have:

-
Experience with knowledge graphs, semantic search, or graph databases (e.g., Neo4j, Amazon Neptune, or similar).

-
Background in real-time inference or low-latency serving requirements.

-
Familiarity with agentic AI systems, autonomous agents, or multi-step reasoning pipelines.

-
Experience with enterprise data infrastructure, data pipelines, or data integration platforms.

Location & Work Arrangement

This is a full-time, on-site role based in San Mateo, CA . On-site collaboration is an important part of how this small, fast-moving team operates. Visa sponsorship is not available for this position.
Why Join

-
Early-stage opportunity with significant ownership and impact โ€” you'll shape foundational infrastructure decisions.

-
Work on genuinely hard distributed systems problems in a production AI context.

-
Small, experienced team with deep ML and enterprise engineering backgrounds.

-
Well-funded at the seed stage with strong institutional backing and a clear enterprise customer focus.

Originally posted on Himalayas

โ† All remote jobs