ML Infrastructure Engineer
About the Role
We're a seed-stage enterprise AI infrastructure company building the context layer that makes AI agents reliable, accurate, and secure for critical business operations โ including highly regulated industries like insurance, banking, asset management, and healthcare. Our platform automatically constructs a governed, real-time domain model across all enterprise data, enabling AI agents to make confident, auditable decisions in production.
As an ML Infrastructure Engineer , you'll own the systems that keep our agents running reliably and fast at scale. This is a hands-on production engineering role โ focused on real-world impact, not research. You'll design, build, and scale our inference and model-serving infrastructure as concurrency and customer demands grow.
What You'll Do
-
Own inference and model-serving infrastructure end to end โ from initial design through production deployment and ongoing scaling.
-
Build and scale systems that enable AI agents to run reliably and efficiently under high and increasing concurrency.
-
Identify and resolve infrastructure bottlenecks in collaboration with ML and platform engineering teams.
-
Optimize systems for latency, throughput, and reliability across cloud-hosted production environments.
-
Drive observability, monitoring, and debugging practices across our production ML stack.
What We're Looking For
Dealbreakers โ all required:
-
5+ years of hands-on experience building and operating machine learning inference systems, model-serving platforms, or ML infrastructure in production environments.
-
Demonstrated experience designing and scaling inference-serving infrastructure using tools such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom systems.
-
Proven ability to optimize production ML systems for latency, throughput, and reliability at scale.
Required skills:
-
Experience with containerization and orchestration (Docker, Kubernetes) for deploying and scaling ML workloads.
-
Background in distributed systems that handle high concurrency and dynamic resource allocation under load.
-
Proficiency with monitoring and observability tooling โ e.g., Prometheus, Grafana, ELK stack, distributed tracing.
-
Experience deploying and managing ML systems on cloud platforms (AWS, GCP, or Azure).
-
Proficiency in at least one systems or backend language: Python, Go, Rust, C++, or Java.
Nice to have:
-
Experience with knowledge graphs, semantic search, or graph databases (e.g., Neo4j, Amazon Neptune, or similar).
-
Background in real-time inference or low-latency serving requirements.
-
Familiarity with agentic AI systems, autonomous agents, or multi-step reasoning pipelines.
-
Experience with enterprise data infrastructure, data pipelines, or data integration platforms.
Location & Work Arrangement
This is a full-time, on-site role based in San Mateo, CA . On-site collaboration is an important part of how this small, fast-moving team operates. Visa sponsorship is not available for this position.
Why Join
-
Early-stage opportunity with significant ownership and impact โ you'll shape foundational infrastructure decisions.
-
Work on genuinely hard distributed systems problems in a production AI context.
-
Small, experienced team with deep ML and enterprise engineering backgrounds.
-
Well-funded at the seed stage with strong institutional backing and a clear enterprise customer focus.
Originally posted on Himalayas