Senior Cloud Infrastructure and DevOps Solutions Architect

🏢 NVIDIA · all 671 jobs
📍 France
📅 Posted Sep 20, 2026 · via Himalayas
🏷 Solutions Architect, Cloud Infrastructure, DevOps Engineer, Hpc Infrastructure, Technical Architecture, Senior Cloud Solutions Architect +7 more
Apply on original site ↗

NVIDIA is looking for a Senior Cloud Infrastructure and DevOps Solutions Architect to join its NVIDIA Infrastructure Specialist Team. Academic and commercial organizations around the world are using NVIDIA products to redefine deep learning and data analytics, and to power next-generation data centers. Join the team building and advising on many of the largest and fastest AI/HPC systems in the world!

We are looking for someone who combines deep technical expertise with strong consulting and communication skills. This role will engage directly with customers, partners, and cross-functional teams to assess, architect, and guide the implementation of large-scale infrastructure projects. The scope spans system architecture, Kubernetes-based platforms, and automation—serving as both a trusted advisor and a hands-on technical leader. You will sit at the centre of NVIDIA 's Cloud Partner (NCP) operating model, covering the full Day 1 to Day 2 lifecycle: taking a GPU cluster from hardware handover, through full-solution validation, to a production-stable platform running at maximum goodput. NCP estates are open-source-first and heterogeneous—upstream Kubernetes, KubeVirt, Slurm, Prometheus/Grafana, Cumulus/SONiC and a long tail of ISV software—so this role is deliberately tool-agnostic: you will meet each partner on the stack they actually run rather than on a single proprietary product.
What You’ll Be Doing:

-
Own full-solution validation on the partner software stack—the layer above hardware validation—including cluster-wide stability testing, real training-workload acceptance, and multi-day, multi-rack burn-in against agreed MTBI and goodput targets.

-
Minimise the time from cluster handover to first production workload, working across hardware bring-up, managed-service intake and the partner's own operations teams to remove duplicated validation and handover friction.

-
Own Day 2 production stability at fleet scale: monitoring, logging and workload orchestration, fault detection and remediation, preventive maintenance, and proactive firmware and field-notice rollout campaigns.

-
Assess customer environments and operate heterogeneous open platforms—upstream Kubernetes, KubeVirt, Slurm and GPU-aware schedulers—integrated with enterprise-grade networking and storage, and enable third-party ISV workloads on top of them.

-
Provide consultative guidance and hands-on troubleshooting across the full stack—bare metal, operating system, software stack, container platform, networking and storage—and support R&D, POCs and POVs validating new features, architectures and upgrade approaches.

-
Act as the technical leader for assigned accounts: run structured knowledge transfer and enablement, and produce runbooks, onboarding materials and best-practice guides so partner teams can operate advanced configurations independently.

What We Need to See:

-
BS/MS/PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields, or equivalent experience.

-
8+ years in managing scalable cloud environments and automation engineering roles.

-
Cloud, HPC & GPU Expertise: Proven understanding of networking fundamentals and data centre architectures, with hands-on experience managing HPC/AI clusters and NVIDIA GPU-accelerated infrastructure—deployment, driver and CUDA toolkit management, optimisation, workload profiling and troubleshooting across CPUs, GPUs and high-speed interconnects.

-
Kubernetes & AI/ML Workloads: Extensive background with Kubernetes for container orchestration, resource scheduling and scaling in GPU-accelerated and HPC environments, including scheduler internals, batch schedulers such as Slurm, and mixed bare-metal/virtualised (e.g. KubeVirt) multi-tenant estates.

-
Linux & Storage Systems: Deep knowledge of Linux (RedHat, Ubuntu), OS-level security, and protocols. Experience with storage solutions such as Lustre, GPFS, ZFS, XFS, and emerging Kubernetes storage technologies.

-
Au

Flights + hotels

This role requires you to be in France. If that means relocating or flying in, it is worth checking fares before you commit to a start date.

Compare flights and hotels →

← All remote jobs

Comparing DevOps Engineer pay and openings — the live median is $119k?All remote DevOps Engineer jobs →DevOps Engineer salary data →
Want more like this? Browse every live remote developer role.All remote developer jobs →
Get new developer jobs by email
Daily email, only when there's something new. One click to stop.

Get remote developer jobs like this by email

10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.

Similar for you