Solution Architect (Storage) (Telecommuter, US)
About the role
Own the architecture of high-performance storage and data platforms supporting AI training, inference, RAG, checkpointing, and large-scale data pipelines, ensuring data can be delivered to GPU infrastructure at the required performance and scale. This role is a customer-facing technical leader supporting the sales organization throughout discovery, solution development, technical validation, and transition to deployment. The architect works as part of the broader AI infrastructure architecture team and collaborates closely with the other specialist domains to deliver an integrated end-to-end solution.
What you'll be doing
- Lead technical discovery and storage architecture for strategic AI Data Center opportunities.
• Design high-performance storage architectures for AI training, inference, checkpointing, RAG, and data-intensive workloads.
• Architect parallel and scale-out file systems, object storage, block storage, local NVMe, and data movement solutions based on workload requirements.
• Translate customer requirements into capacity, throughput, IOPS, latency, metadata, resiliency, and bandwidth-per-GPU targets.
• Size storage platforms based on GPU cluster scale, workload characteristics, dataset growth, checkpoint behavior, and performance requirements.
• Design compute-to-storage connectivity and collaborate on storage network architecture, RDMA, Ethernet/InfiniBand, and data-path optimization.
• Develop storage reference architectures, BOMs, sizing models, performance assumptions, and technical standards.
• Evaluate AI-focused storage platforms, technology roadmaps, interoperability, data services, and lifecycle requirements.
• Lead storage benchmarking, validation, performance tuning, troubleshooting, and production readiness activities.
• Partner with GPU / Compute and Network Solution Architects to eliminate data-path bottlenecks and optimize end-to-end cluster performance.
• Support RFQs/RFPs, proposals, customer workshops, proofs of concept, technical presentations, and architecture reviews.
• Work with Services Architecture, integration, OEMs, Professional Services, and delivery teams to ensure storage solutions can be deployed and operated at scale.
What you have
- 7+ years of experience in enterprise storage, high-performance storage, distributed systems, HPC, solution architecture, or related technical roles.
• Deep expertise in high-performance storage, distributed storage, HPC storage, or AI data platforms.
• Strong knowledge of parallel file systems, scale-out NAS, object storage, block storage, NVMe, caching, metadata, and data movement.
• Experience translating workload requirements into capacity and performance sizing across throughput, IOPS, latency, metadata, and resiliency.
• Experience supporting AI/ML, HPC, hyperscale, cloud, or other data-intensive infrastructure environments.
• Understanding of storage networking, RDMA, Ethernet/InfiniBand, and the interaction between storage, GPU compute, and network fabrics.
• Strong performance engineering and troubleshooting skills across storage systems and end-to-end data paths.
• Strong customer-facing architecture, documentation, workshop, and presentation skills.
• Ability to work across customers, OEMs, compute/network architects, Professional Services, integration, and delivery teams.
• VAST Data, WEKA, DDN, Pure Storage, Dell, NetApp or comparable high-performance storage platforms
• Lustre, IBM Storage Scale/GPFS, Ceph or other distributed/parallel storage technologies
• AI training, inference, checkpointing, RAG, and large-scale data pipeline architectures
• NVMe and high-performance storage networking
• Storage benchmarking, performance modeling, telemetry, and optimization
• AI CSP, hyperscale, NCP / neocloud, HPC, or large cloud infrastructure environments
• Technical Architecture – Brings deep domain expertise and translates customer requirements into scalable, supportable solutions.
• AI Infrastructure Knowledge – Understands