Technical Account Manager
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI cloud, join us.
*Note: This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.
The Technical Account Manager owns the technical health of the post-sales relationship for Lambda’s public cloud accounts, spanning AI-native startups, Enterprises, and Fortune 500 companies. Where the Customer Success Manager owns the commercial health of an account, you own its technical health: the customer’s workloads run well, the architecture is right, the SLA story is defensible, and technical risk is found and retired before it threatens revenue. Solutions Engineering carries the account through pre-sales and hypercare; at handoff, you take ownership of the technical relationship for the life of the contract.
This is a hands-on role, not a coordination role. You will understand what customers are actually building (training runs, fine-tuning pipelines, inference services) deeply enough to lead joint POC sessions, design and defend architectures, validate SLA events at the root-cause level, and build the tooling that makes account health measurable. You will be the customer’s most credible technical advocate inside Lambda and Lambda’s most trusted technical voice inside the account.
What You’ll Do
- Own the technical health of your accounts. Take the technical handoff from Solutions Engineering at the end of hypercare and own the account’s technical outcomes through steady state, expansion, and renewal. Know the state of every cluster and workload you are accountable for, and keep your commercial counterparts ahead of technical risk.
- Understand customer AI use cases end to end. Map what it means for each customer to train, fine-tune, and serve models on Lambda: frameworks, schedulers, parallelism strategy, data paths, and performance baselines. Build the customer user journey and convert it into value-add opportunities across the platform, documentation, and escalation routing.
- Lead joint customer POC sessions. Define success criteria with the customer before a node is provisioned: acceptance thresholds, benchmarks, timelines. Own the execution plan, coordinate capacity and provisioning, run or oversee the tests, and drive the POC to a clear verdict: win the workload, close the gap through product, or qualify out.
- Lead customer architecture designs. Produce and defend reference architectures spanning compute, networking, storage, connectivity, and scheduler integration (Slurm, Kubernetes). Make support boundaries explicit: what is managed and what is not. Own the design as it evolves after handoff, pulling in engineering domain experts with specific, well-framed questions.
- Own SLA and reliability engineering. Build and own the canonical methodology for uptime, downtime, and credit calculation. Validate breach events at the technical level, down to the specific Ethernet or InfiniBand failure, and arm CSMs and leadership with defensible numbers. Partner with product to standardize SLA language and structure across 1CC, on-demand, and reserved offerings.
- Build the tooling that makes accounts measurable. Own the technical data surfaces for customer health end to end: dashboards, telemetry and uptime history views, health scoring, and churn early-warning signals. Scope, build, and drive adoption. Replace “escalate to engineering to answer a basic question” with self-serve data for the whole GTM org.
- Direct technical escalations and incidents. Serve as the technical lead during high-severity events on your accounts: drive root cause, hold the quality b