Senior Principal AI Engineer
A Moving Experience.
What You Will Work On
-
Design andoperatedistributed training systems for large neural networks(autoregressive, diffusion,State Space Modelsetc.)across GPU clusters
-
Optimisemulti‑node, multi‑GPU execution tomaximizethroughput andutilization
-
Diagnose&resolve bottlenecks across compute, memory, and network
-
Improve training stability and fault tolerance at scale
-
Partner with research and applied ML teams to productionizelarge‑modeltraining pipelines
Core Responsibilities
- Distributed Training Infrastructure
-
Build andoptimizeGPU cluster orchestration using:
- Slurm
- Kubernetes
- Ray
- RunAI
-
Ensure efficient scheduling, isolation, and fairness across training workloads
- Communication & Networking
-
Optimizeand debug distributed communication using:
- NCCL
- RDMA
- InfiniBand
- NVLink
-
Minimizenetworking bottlenecks that dominateend‑to‑endtraining time
- Training Frameworks
- Scale large-model training using:
- PyTorchDistributed
- Megatron‑LM
- DeepSpeed
-
Ownmulti‑nodelaunch configurations, failure recovery, and performance tuning
- Memory & Performance Optimization
-
Apply advanced memory optimization techniques:
- Activation checkpointing
- ZeRO(Stage 1–3) and offload strategies
-
Balance compute, memory, and communication to push model size and batch scale
What Success Looks Like
-
GPUutilizationconsistently stays high (>80–90%)
-
Training scales cleanly from single node to dozens or hundreds of GPUs
-
Communication overhead is minimized and predictable
-
Large training jobs run stably for days or weeks without failure
-
New models can be trained faster, larger, and more reliably than before
Required Experience & Skills
- Strongly Required
-
Deephands‑onexperience with distributed systems or ML systems
-
Experience runninglarge‑scaleworkloads on GPU clusters
-
Production experience withPyTorchdistributed training
-
Strong understanding of parallelism strategies (data, tensor, pipeline parallelism)
-
Low‑levelunderstanding of GPU communication and networking
- Critical Technical Skills
-
GPU orchestration:Slurm, Kubernetes, Ray,RunAI
-
Communication libraries: NCCL, RDMA, InfiniBand,NVLink
-
Training frameworks:PyTorchDistributed,Megatron‑LM,DeepSpeed
-
Memoryoptimisation: activation checkpointing,ZeROoffload techniques
Common ProblemsYou’llBe Solving
- Many teams fail at scale because:
-
GPUutilizationis low despite large clusters
-
Networking and communication dominate training time
-
Training jobs crash or become unstable at large scale
-
You will be explicitly focused oneliminatingthese failure modes.
Ideal Background
-
This role is a strong fit for individuals who have worked as:
- ML Systems Engineer
- Distributed Systems Engineer
- AI Infrastructure Engineer
- HPC Engineer transitioning into ML
-
Experience working with large language models or foundation models is a strong plus, but deep systemsexpertiseis valued over pure model architecture experience.
Why This Role Matters
Without robust distributed training infrastructure, progress on largemodelsstalls. This role directly enables:
- Larger models
- Faster iteration cycles
-
More reliable research-to-production pipelines
You will be building the foundation that makeslarge‑scaleAI possible.
Cerence Inc. (Nasdaq: CRNC and ) is the global industry leader in creating unique, moving experiences for the automotive world. Spun out from Nuance in October 2019, Cerence is a new, independent company that has quickly gained traction as a leader in the automotive voice assistant space, working with all of the world’s leading automakers – from Ford and Fiat Chrysler to Daimler, Audi and BMW to Geely and SAIC – to transform how a car feels, responds and learns. Its track record is built on more than 20 years of industry expe