AI Solution Architect

🏢 uvation · all 25 jobs
📍 India
📅 Posted Sep 20, 2026 · via Himalayas
🏷 AI Infrastructure Architect, Gpu Solutions Architect, Infrastructure Architecture, AI Platform Engineering, Cloud Infrastructure Architecture, AI Solution Architect +8 more
Apply on original site ↗

Job Overview

We are seeking an experienced AI Solution Architect to design and lead end-to-end enterprise AI Factory and GPU infrastructure solutions spanning compute, high-performance networking, storage, Kubernetes, cloud, and AI/ML platforms. The role requires strong expertise in NVIDIA GPU technologies, AI workloads, scalable infrastructure architecture, security, observability, performance engineering, and capacity planning .

Key Responsibilities

- Own end-to-end architecture for AI Factory and enterprise AI solutions from requirements through production readiness.

- Assess AI/ML workload requirements for training, fine-tuning, inference, batch processing, and high-performance computing.

- Design GPU compute architectures including NVIDIA HGX/DGX/OEM platforms, multi-GPU systems, NVLink/NVSwitch, and GPU resource allocation.

- Design high-performance AI networking using 100/200/400/800G Ethernet, EVPN/VXLAN, and leaf-spine architectures.

- Design AI storage and data architectures using object storage, parallel file systems like Ceph, WEKA, , or equivalent platforms.

- Define AI platform architecture across Kubernetes, HPC, container runtimes, model-serving platforms, and enterprise AI frameworks.

- Establish architecture standards for security, identity, tenant isolation, data protection, observability, disaster recovery, and operational resilience.

- Develop reference architectures, high-level/low-level designs, capacity models, bills of materials, technology evaluations, and implementation roadmaps.

- Lead technical evaluations, proof-of-concepts, vendor assessments, and architecture review boards.

- Collaborate with infrastructure, network, security, storage, cloud, data, application, and operations teams.

- Define performance, availability, scalability, security, and cost objectives and validate architecture against measurable acceptance criteria.

- Provide technical leadership during deployment, migration, integration, troubleshooting, and production transition.

• Required Technical Skills

AI / ML Architecture

• NVIDIA AI Enterprise, NGC, CUDA, NCCL, DCGM, GPU Operator and AI platform ecosystem.

• PyTorch, TensorFlow, JAX and operational understanding of training and inference workloads.

• GPU scheduling, multi-tenancy, MIG/vGPU, GPU utilization and workload placement.

• LLM, generative AI, RAG, fine-tuning, model serving and inference architecture.

GPU & AI Factory Infrastructure

• NVIDIA A100/H100/H200/B200 or equivalent GPU platforms; familiarity with next-generation systems.

• NVLink, NVSwitch, PCIe topology and multi-GPU performance architecture.

• DGX/HGX/OEM GPU server architecture and lifecycle management.

• AI Factory capacity planning, rack density, power, cooling, commissioning and lifecycle strategy.

High-Performance Networking

• 100/200/400/800G Ethernet, InfiniBand, RoCEv2 and RDMA, Netris

• NVIDIA ConnectX/SuperNIC, Spectrum/Spectrum-X, Quantum and BlueField DPU technologies.

• BGP, EVPN/VXLAN, VRF, ECMP, VLAN, MTU, PFC, ECN, QoS and congestion management.

• GPU east-west traffic, GPUDirect RDMA and network performance troubleshooting.

AI Storage & Data Architecture

• Parallel file systems, object storage, NFS, NVMe/NVMe-oF and high-throughput data pipelines.

• Ceph, WEKA, VAST, Dell PowerScale, Pure FlashBlade, NetApp or equivalent technologies.

• Data lake/lakehouse concepts, metadata, lineage, data movement and data lifecycle.

• GPUDirect Storage and storage/network performance optimization.

AI Platform & Orchestration

• Kubernetes, GPU Operator, container runtimes and Kubernetes GPU scheduling.

• HPC or other equivalent workload schedulers.

• Model serving/inference platforms and MLOps platform architecture.

• API gateways, service discovery, secrets management and platform integration.

Cloud & Hybrid Architecture

• AWS and/or Azure AI infrastructure and security services.

• Hybrid cloud connectivity, IAM, private networking, cloud storage and work

Flights + hotels

This role requires you to be in India. If that means relocating or flying in, it is worth checking fares before you commit to a start date.

Compare flights and hotels →

← All remote jobs

Want more like this? Browse every live remote data science role.All remote data science jobs →
Get new data science jobs by email
Daily email, only when there's something new. One click to stop.

Get remote data science jobs like this by email

10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.

Similar for you