Technical Support Engineer (GPU Clusters) - US Weekends

🏢 Together AI · all Together AI jobs
📍 United States
💰 USD 160,000 - 230,000 / annual
📅 Posted 2026-08-06 · via Himalayas
🏷 Technical-Support-Engineer,Customer-Success,Site-Reliability-Engineering,AI-Infrastructure-Support,GPU-Cluster-Engineering,Senior-Technical-Support-Engineer,Support-Engineer
Apply on original site ↗

About the role

As a Technical Support Engineer at a pioneering AI company, you'll be the first line of defense to support customers as they build out training, fine tuning, and inference solutions with Together AI . You'll dive deep into complex technical challenges, providing swift and effective solutions while serving as a product expert. As a part of the Customer Experience organization, you will collaborate closely with product and sales, driving continuous improvement of our offerings. This is an exciting opportunity for a deeply technical professional passionate about AI and customer success to make a significant impact in a fast-paced, innovative environment.
Required hours

- This is a fulltime position working US daytime hours. The role will work both weekend days (Saturday and Sunday) as well as two additional weekdays.

- This is a 4-day shift, 10 hours per day, with 2 additional hours of on-call coverage on Saturdays and Sundays.

- The role would start as a Monday to Friday role for the first few months to allow for ramping up and learning from teammates. After being considered fully ramped, the role would transition to the 4-day weekend shift.

Responsibilities

- Engage directly with customers to tackle and resolve complex technical challenges involving our cutting-edge Kubernetes GPU clusters; ensure swift and effective solutions every time.

- Act as a customer facing SRE to ensure our customer’s Kubernetes clusters remain healthy and stable

- Become a product expert in our GPU Cluster service, serving as the last line of technical defense before issues are escalated to Engineering and Product teams.

- Monitor GPU cluster health and proactively communicate hardware issues to customers (thermal throttling, BMC failures, missing GPUs, and NVLink/InfiniBand degradation) with clear remediation steps

- Operate and maintain production infrastructure for enterprise GPU customers, including fleet rebalancing, Slurm cluster maintenance, node repair/migration, and Kubernetes-based workload management

- Investigate and resolve storage and networking issues such as Weka filesystem degradation, InfiniBand link failures, and bandwidth anomalies on bare-metal and VM environments

- Collaborate seamlessly across Engineering, Research, and Product teams to address customer concerns; collaborate with senior leaders both internally and externally to ensure the highest levels of customer satisfaction.

- Transform customer insights into action by identifying patterns in support cases and working with Engineering and Go-To-Market teams to drive Together’s roadmap (e.g., future models to support)

- Maintain detailed documentation of system configurations, procedures, troubleshooting guides, and FAQs to facilitate knowledge sharing with team and customers.

- Be flexible in providing support coverage during holidays, nights and weekends as required by business needs to ensure consistent and reliable service for our customers.

Requirements

- 3+ years of experience in a customer-facing technical role with at least 1 year in a support function for an AI service or supporting a mission-critical API in SaaS

- Experience as an SRE or DevOps engineer working with Kubernetes

- Strong technical background, with knowledge of AI, ML, GPU technologies and their integration into high-performance computing (HPC) environments.

- Advanced knowledge with infrastructure services (e.g., Kubernetes, SLURM), infrastructure as code solutions (e.g., Ansible) high-performance network fabrics, NFS-based storage management, container infrastructure, and scripting and programming languages.

- Experience with HPC/Slurm cluster environments — node draining, job scheduling, maintenance workflows

- Familiarity with high-speed networking concepts — InfiniBand, RDMA, network interface diagnostics

- Experience with distributed storage systems (e.g., Weka, NFS) and troubleshooting I/O and bandwidth issues

- Foundational understanding in the inst

← All remote jobs