Technical Support Engineer (Inference) - US Weekends
About the role
As a Technical Support Engineer at a pioneering AI company, you'll be the first line of defense to support customers as they build out training, fine tuning, and inference solutions with Together AI . You'll dive deep into complex technical challenges, providing swift and effective solutions while serving as a product expert. As a part of the Customer Experience organization, you will collaborate closely with product and sales, driving continuous improvement of our offerings. This is an exciting opportunity for a deeply technical professional passionate about AI and customer success to make a significant impact in a fast-paced, innovative environment.
Required hours
- This is a fulltime position working US daytime hours. The role will work both weekend days (Saturday and Sunday) as well as two additional weekdays.
- This is a 4-day shift, 10 hours per day, with 2 additional hours of on-call coverage on Saturdays and Sundays.
- The role would start as a Monday to Friday role for the first few months to allow for ramping up and learning from teammates. After being considered fully ramped, the role would transition to the 4-day weekend shift.
Responsibilities
- Engage directly with customers to tackle and resolve complex technical challenges involving our cutting-edge GPU clusters and our inference and fine-tuning services; ensure swift and effective solutions every time.
- Act as a customer facing SRE to ensure our customer’s Inference endpoints (running on Kubernetes) remain healthy, stable, and performant
- Become a product expert in all of our Gen AI solutions, serving as the last line of technical defense before issues are escalated to Engineering and Product teams.
- Assist with hardware and platform migrations by validating system health and traffic routing. Monitor dashboards to detect anomalies and escalate with data-backed analysis
- Manage customer-facing communications during incidents and degradations; translate deep technical findings (latency regressions, provider issues, network reachability drops) into clear, evidence-backed updates without exposing platform internals
- Contribute infrastructure changes for model deployment, capacity rebalancing, and cluster configuration. You will execute infrastructure changes via pull requests (infra-as-code) for tasks such as endpoint configuration, model bringup/bringdown, and capacity scaling
- Flag engine-level bugs with logs and reproduction steps for engineering
- Collaborate seamlessly across Engineering, Research, and Product teams to address customer concerns; collaborate with senior leaders both internally and externally to ensure the highest levels of customer satisfaction.
- Transform customer insights into action by identifying patterns in support cases and working with Engineering and Go-To-Market teams to drive Together’s roadmap (e.g., future models to support)
- Maintain detailed documentation of system configurations, procedures, troubleshooting guides, and FAQs to facilitate knowledge sharing with team and customers.
- Be flexible in providing support coverage during holidays, nights and weekends as required by business needs to ensure consistent and reliable service for our customers.
Requirements
- 6+ years of experience in a customer-facing technical role, SRE, DevOps, or infrastructure engineering, with at least 1 year in a support role for an AI service
- Experience as an SRE or DevOps engineer working with Kubernetes
- Strong technical background, with knowledge of AI, ML, GPU technologies and their integration into high-performance computing (HPC) environments.
- Advanced, production-level experience with infrastructure services (e.g., Kubernetes, SLURM), infrastructure as code solutions (e.g., Ansible) high-performance network fabrics, NFS-based storage management, and container infrastructure
- Familiarity with operating storage systems in HPC environments such as Vast and Weka
- Proven ability to diagnos