Sr. Solution Engineer/Architect (GPU Infrastructure)
๐ข Nebius ยท all Nebius jobs
๐ United States
๐ฐ USD 180,000 - 220,000 / annual
๐
Posted 2026-08-23 ยท via Himalayas
๐ท Solutions-Engineering,Solutions-Architect,Customer-Engineering,Hardware-Infrastructure,Cloud-Infrastructure,GPU-Infrastructure-Engineering,Technical-Account-Management,GPU-Solutions-Architect,GPU-Infrastructure-Engineer,GPU-Architecture-Engineer,GPU-Systems-Engineer,GPU-Cloud-Engineer,Senior-AI-Infrastructure-Engineer,GPU-Computing-Engineer,Senior-Solutions-Engineer
Apply on original site โAbout Nebius :
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
The role
Nebius is looking for a Sr. Solution Engineer, you will act as the primary technical partner for customers deploying and operating GPU clusters and AI infrastructure on Nebius . You will bridge customer requirements with internal engineering capabilities, ensuring successful deployment, stable operations, and ongoing optimization of complex, high-performance environments.
Your responsibilities will include:
- Acting as the main technical interface for customers running workloads on Nebius GPU infrastructure.
- Supporting customers in deploying, configuring, and tuning GPU-based environments for performance and reliability.
- Investigating and resolving complex issues spanning hardware, networking, operating systems, and cluster-level behavior.
- Partnering with internal teams (data center operations, networking, platform engineering) to coordinate and drive resolution of customer-impacting issues.
- Converting customer requirements into practical architectures, configurations, and execution plans.
- Identifying opportunities to improve system performance, stability, and overall customer experience.
- Developing and maintaining technical documentation, including solution patterns, troubleshooting guides, and operational best practices.
- Contributing to continuous improvement by surfacing recurring issues, gaps, and optimization opportunities to internal teams.
We expect you to have:
- Experience in a customer-facing technical role (e.g., solutions engineer, support engineer, technical account manager, or similar).
- Strong understanding of GPU infrastructure, including NVIDIA-based systems, multi-node environments, and performance considerations.
- Hands-on experience with Linux systems and system-level troubleshooting.
- Familiarity with large-scale compute environments such as GPU clusters, AI infrastructure, or supercomputing systems.
- Ability to diagnose issues across hardware, networking, and software layers.
- Strong analytical and problem-solving skills, with the ability to operate effectively in high-pressure situations.
- Excellent communication skills, with the ability to explain complex technical concepts clearly to customers.
- A proactive, ownership-driven approach with a strong focus on customer success.
It would be an added bonus if you have:
- Experience working with NVIDIA Grace Blackwell or similar next-generation GPU platforms.
- Experience with cluster validation, benchmarking, or performance testing tools (e.g., HPL, NCCL)
Key employee benefits:
- Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan: up to 4% company match with immediate vesting.
- Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.
- Remote work reimbursement: up to $85/month for mobile and internet.
- Disability & life insurance: company-paid short-term, long-term and life insurance coverage.
Compensation
We offer competitive salaries ranging from $180K to $220K OTE, which includes base salary and performance bonus. Equity in the form of RSUs may be available at certain salary grades.
Join Nebius Today!
Benefits & Perks:
- Competitive compensati