Staff ML Infrastructure Engineer - Embodied AI

🏢 General Motors · all General Motors jobs
📍 United States
💰 USD 189,300 - 290,700 / annual
📅 Posted 2026-07-02 · via Himalayas
🏷 Software-Engineer,ML-Infrastructure-Engineering,ML-Platform-Engineering,Autonomous-Vehicle-Engineering,AI-Infrastructure,AI-ML-Infrastructure-Engineer,ML-Infrastructure-Engineer,AI-Infrastructure-Engineer,Machine-Learning-Infrastructure-Engineer,Staff-ML-Engineer,AI-Data-Infrastructure-Engineer,Staff-Applied-AI-Engineer,Staff-AI-Engineer,Senior-AI-ML-Operations-Engineer
Apply on original site ↗

Job Description

At General Motors , our product teams are redefining mobility. Through a human-centered design process, we create vehicles and experiences that are designed not just to be seen, but to be felt. We’re turning today’s impossible into tomorrow’s standard —from breakthrough hardware and battery systems to intuitive design, intelligent software, and next-generation safety and entertainment features.

Every day, our products move millions of people as we aim to make driving safer, smarter, and more connected, shaping the future of transportation on a global scale.
Role:

Are you passionate about accelerating the future of autonomous driving? Join the Embodied AI team at General Motors . Our team is developing and deploying machine learning solutions that support safe and reliable autonomous vehicle behavior across real-world scenarios.

As a Staff ML Infra Engineer, you willdrive the development of core systems that enable rapiddataset generation,training, evaluation, and iteration of our most advanced Autonomous Driving models.Fromenablinglargefoundational drivingmodels todistilling multi-stage production deployedmodels,yourgoalwill be todramatically accelerate the machine learning development cyclefrom one modeling hypothesis to next.

You will deliver model training pipelinesthat are performant, easy to use, and exceptionally reliable.Your success will be measured by the velocity and impact of the ML models that rely on the scalable, intuitive, and high‑performance training platforms you help create.

What you'll do:
-
Lead the design, implementation, and deployment of scalable platforms and tools that drive machine learning model training and evaluation workflows across GM.

-
Own complex technical projects end-to-end, making key architectural decisions and technical trade-offs. You will be a core contributor to team planning, design reviews, and code quality.

-
Take a holistic view of projects, considering their impact across multiple teams, and across a longer timeline.

-
Proactively drive technical prioritization. Collaborate closely with partner teams to ensure maximum benefit from the systems we build.

-
Help shape our team through technical interviewing with high, well-calibrated standards, and play an essential role in recruiting.

-
Mentor and onboard junior engineers and interns, helping them grow their careers.

What you'll bring:
-
5+ years of experience building large-scale distributed systems, applications, or advanced ML systems‑scale distributed systems, applications, or advanced ML systems

-
Proventrack recordof designing robust frameworks with high-quality, durable APIs.

-
Deep understanding of machine learning algorithms with hands‑on application

-
Expertisein building reliable, high-performance, and cost-efficient systems on modern cloud infrastructure‑performance

-
End-to-end experience across the ML development lifecycle, includingMLOpspractices

-
Strong cross functional collaboration skills across teams and organizations

-
Exceptional coding skills in Python or C++

-
Strong interest in autonomous driving and its transformative potential

-
BS, MS, or PhD in Computer Science, Mathematics, or equivalent practical experience

Nice to have:
-
Experience with distributed training methodologies

-
Experience scaling ML training across large GPU/CPU clusters or other accelerators

-
Familiarity with deep learning frameworks (e.g.,PyTorch, TensorFlow)

-
Experience with performance profiling andstate-of-the-arttraining optimization techniques, including their impact on model performance ‑of‑the‑art training optimization techniques, including their impact on convergence.

-
Experience with advanced build systems (e.g., Bazel, Buck, Blaze,CMake)

-
Proficiencywith containerization and orchestration technologies (e.g., Docker, Kubernetes)

Remote/Hybrid: This role is categorized as fully remote or hybrid.

Compensation: The compensation infor

← All remote jobs