Senior Principal AI Engineer

🏢 Cerence · all Cerence jobs
📍 United States
📅 Posted 2026-08-08 · via Himalayas
🏷 AI-Engineering,ML-Infrastructure-Engineering,Distributed-Systems-Engineering,AI-Infrastructure-Engineer,ML-Systems-Engineer,Principal-AI-Engineer,Senior-AI-Engineer,AI-Principal-Engineer,Principal-AI-ML-Engineer,Senior-AI-ML-Engineer
Apply on original site ↗

A Moving Experience.

What You Will Work On

-
Design andoperatedistributed training systems for large neural networks(autoregressive, diffusion,State Space Modelsetc.)across GPU clusters

-
Optimisemulti‑node, multi‑GPU execution tomaximizethroughput andutilization

-
Diagnose&resolve bottlenecks across compute, memory, and network

-
Improve training stability and fault tolerance at scale

-
Partner with research and applied ML teams to productionizelarge‑modeltraining pipelines

Core Responsibilities

- Distributed Training Infrastructure

-
Build andoptimizeGPU cluster orchestration using:

- Slurm

- Kubernetes

- Ray

- RunAI

-
Ensure efficient scheduling, isolation, and fairness across training workloads

- Communication & Networking

-
Optimizeand debug distributed communication using:

- NCCL

- RDMA

- InfiniBand

- NVLink

-
Minimizenetworking bottlenecks that dominateend‑to‑endtraining time

- Training Frameworks

- Scale large-model training using:

- PyTorchDistributed

- Megatron‑LM

- DeepSpeed

-
Ownmulti‑nodelaunch configurations, failure recovery, and performance tuning

- Memory & Performance Optimization

-
Apply advanced memory optimization techniques:

- Activation checkpointing

- ZeRO(Stage 1–3) and offload strategies

-
Balance compute, memory, and communication to push model size and batch scale

What Success Looks Like

-
GPUutilizationconsistently stays high (>80–90%)

-
Training scales cleanly from single node to dozens or hundreds of GPUs

-
Communication overhead is minimized and predictable

-
Large training jobs run stably for days or weeks without failure

-
New models can be trained faster, larger, and more reliably than before

Required Experience & Skills

- Strongly Required

-
Deephands‑onexperience with distributed systems or ML systems

-
Experience runninglarge‑scaleworkloads on GPU clusters

-
Production experience withPyTorchdistributed training

-
Strong understanding of parallelism strategies (data, tensor, pipeline parallelism)

-
Low‑levelunderstanding of GPU communication and networking

- Critical Technical Skills

-
GPU orchestration:Slurm, Kubernetes, Ray,RunAI

-
Communication libraries: NCCL, RDMA, InfiniBand,NVLink

-
Training frameworks:PyTorchDistributed,Megatron‑LM,DeepSpeed

-
Memoryoptimisation: activation checkpointing,ZeROoffload techniques

Common ProblemsYou’llBe Solving

- Many teams fail at scale because:

-
GPUutilizationis low despite large clusters

-
Networking and communication dominate training time

-
Training jobs crash or become unstable at large scale

-
You will be explicitly focused oneliminatingthese failure modes.

Ideal Background

-
This role is a strong fit for individuals who have worked as:

- ML Systems Engineer

- Distributed Systems Engineer

- AI Infrastructure Engineer

- HPC Engineer transitioning into ML

-
Experience working with large language models or foundation models is a strong plus, but deep systemsexpertiseis valued over pure model architecture experience.

Why This Role Matters

Without robust distributed training infrastructure, progress on largemodelsstalls. This role directly enables:

- Larger models

- Faster iteration cycles

-
More reliable research-to-production pipelines

You will be building the foundation that makeslarge‑scaleAI possible.

Cerence Inc. (Nasdaq: CRNC and ) is the global industry leader in creating unique, moving experiences for the automotive world. Spun out from Nuance in October 2019, Cerence is a new, independent company that has quickly gained traction as a leader in the automotive voice assistant space, working with all of the world’s leading automakers – from Ford and Fiat Chrysler to Daimler, Audi and BMW to Geely and SAIC – to transform how a car feels, responds and learns. Its track record is built on more than 20 years of industry expe

← All remote jobs