Principal Software Engineer, AI Networking

🏢 Nvidia · all 709 jobs
📍 2 Locations
📅 Posted Posted 30+ · via Workday
🏷 Remote
Apply on original site ↗

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

Join NVIDIA, where the future is defined by our innovative advances in AI, computer graphics, and accelerated computing. As a Principal Software Engineer, you will lead the transformation of AI networking systems. You will apply your deep expertise to manage complex customer engagements and help develop our product and architecture direction. This role offers an outstanding opportunity to influence NVIDIA's networking technologies and make a significant impact on the industry!

What you'll be doing:

Lead the technical strategy for AI Factory networking deployments at strategic customers, including conducting architecture reviews, risk assessments, and crafting multi-phase execution plans.

Serve as the principal-level technical authority for embedded networking products like BlueField and ConnectX. This role also covers the surrounding technology ecosystem, including DOCA, RDMA, RoCE, and Infiniband.
-
Lead deep technical engagements with hyperscalers and AI Factory customers, involving design-in, coding, bring-up, performance tuning, failure analysis, and production hardening.

-
Partner with internal engineering, product, and architecture teams to transform customer needs into product features, reference architectures, tooling, and guidelines.

-
Drive performance, reliability, and debuggability improvements across customer stacks and translate findings into actionable product, firmware, and software roadmap items.

What we need to see:
-
BS/MS/PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.

-
15+ years of relevant industry experience, including technical leadership across complex systems.

-
Deep knowledge of networking protocols and distributed systems, with a strong understanding of RoCE/InfiniBand, L1–L4 fundamentals, and performance/latency tradeoffs.

-
Proven low-level software expertise with proficiency in C/C++ and comfort debugging across firmware, driver, and user space.

-
Demonstrated experience in high-performance networking and system-level debugging, including packet drops, retransmissions, congestion, QoS, ordering, and buffer management.

-
Excellent interpersonal skills, with the ability to clearly explain complex topics to engineers, PMs, and customer collaborators, and align cross-organizational teams toward a decision.

Ways To Stand Out from the crowd:
-
Prior experience in customer-facing technical leadership at hyperscalers/CSPs/AI factories (or similarly complex production environments).

-
Hands-on expertise with DPDK, DOCA, RDMA verbs, NCCL, CUDA-aware networking, congestion control, and performance tuning at scale.

-
Experience building internal tools, telemetry, and automation that improve triage speed and operational excellence.

-
Demonstrated innovation: patents, publications, hackathons, rapid prototyping, or shipping new architecture/features end-to-end.

-
Experience leading multi-team initiatives across geo/time zones, with clear examples of influence without authority as well as eager and proactive in bringing to bear AI-powered tools to accelerate debugging, documentation, and day-to-day engineering efficiency while maintaining strong engineering judgment.

With competitive salaries and a generous benefits package, we are widely

Flights + hotels

This role requires you to be in 2 Locations. If that means relocating or flying in, it is worth checking fares before you commit to a start date.

Compare flights and hotels →

← All remote jobs

Comparing Software Engineer pay and openings — the live median is $129k?All remote Software Engineer jobs →Software Engineer salary data →
Get new remote jobs like this by email
Daily email, only when there's something new. One click to stop.

Get remote jobs like this by email

10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.

Similar for you