Senior Compiler Engineer

🏢 NVIDIA · all NVIDIA jobs
📍 United States
💰 USD 152,000 - 241,500 / annual
📅 Posted 2026-09-04 · via Himalayas
🏷 Software-Engineer,Language-Design,Compiler-Engineer,Compiler-Optimization-Engineer,Compiler-Developer,Deep-Learning-Compiler-Engineer,Frontend-Compiler-Engineer,JIT-Compiler-Engineer,Compiler-Engineering,GPU-programming,Systems-Programming
Apply on original site ↗

NVIDIA is dedicated to reinvent accelerated computing. Reinvention requires great technology and amazing people. As an NVIDIA N, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

NVIDIA is hiring a Senior Compiler Engineer to join our team driving the next generation of GPU systems programming. We are redefining how developers write high-performance GPU software by bringing the safety, expressiveness, and modern tooling of Rust to native GPU and CUDA development. On this team, you will build cutting-edge compiler pipelines, custom intermediate representation (IR) frameworks, and JIT compilation systems that bridge host and device execution—allowing developers to write memory-safe, high-performance GPU kernels in idiomatic Rust.
What you’ll be doing:

-
Develop Rust-to-GPU Compiler Pipelines: Design, implement, and maintain custom compiler backends (such as rustc codegen backends and proc-macros) and Rust-native intermediate representation (IR) frameworks to compile standard Rust directly to high-performance CUDA PTX and machine code.

-
Build compiler IRs and Optimizers: Work with modern compiler architectures to lower Rust AST and MIR into IRs including MLIR, PTX, and LLVM, including GPU-specific optimizations.

-
Support complex ahead-of-time, just-in-time, and link time optimization workflows: Build state-of-the-art tooling to support users targeting a broad family of NVIDIA GPUs, host platforms, and feature sets.

-
Define Safe Parallel Abstractions: Architect innovative compiler-enforced safety models that extend Rust's ownership, borrowing, and lifetime disciplines across the GPU launch boundary—preventing data races and enforcing memory safety during asynchronous GPU execution.

-
Expose Next-Gen Hardware Features: Implement type-safe, ergonomic device-side abstractions in Rust for low-level GPU hardware primitives, including shared memory, barriers, scoped atomics, Tensor Memory Accelerator (TMA), and warp/cluster-level operations.

-
Build the future: Supporting today’s accelerated computing applications is not sufficient. We have to build composable building blocks for others around us to build the applications of tomorrow.

What we need to see:

-
Bachelor's, Master's, or Ph.D. in Computer Science, Computer Engineering, a related field, or equivalent experience.

-
5+ years of relevant work or research experience in compiler development, language design, or GPU code generation.

-
Deep expertise in the Rust programming language, including a strong grasp of compiler internals (rustc), Rust MIR, procedural macros, and Rust's borrow-checker/lifetime model.

-
Hands-on experience with compiler infrastructures, intermediate representations (IRs), and code generation (such as LLVM IR, MLIR, or custom IR systems).

-
Solid understanding of parallel programming models, GPU architectures, and CUDA programming.

-
Strong software design skills, including debugging, profiling, and benchmarking compilers and GPU kernels.

-
Ability to orchestrate agents for product requirement design, architecture, code development, testing, code review, and issue triage.

-
Ability to work independently, define project goals and scope, and drive complex compiler-engineering efforts from research to production.

Ways to stand out from the crowd:

-
A track record of contributing to the Rust compiler (rustc), Cargo tooling, or open-source Rust-to-GPU projects.

-
Experience building custom compiler front-ends, AST translators, or JIT engines.

-
Familiarity with MLIR or other extensible compiler frameworks.

-
Deep proficiency in low-level GPU programming, including the use of modern hardware features (e.g., Tensor Cores, warp-level shuffles, or asynchronous transfer pipelines).

-
Experience designing Domain-Specific Languages (DSLs) or tile-based programming abstractions for tensor processin

← All remote jobs

Similar for you