Research Analyst - AI System Performance Modelling

🏢 SemiAnalysis · all SemiAnalysis jobs
📍 North Korea,South Korea
📅 Posted 2026-08-03 · via Himalayas
🏷 AI-Research-Analyst,Performance-Modeling,ML-Systems-Analyst,AI-Infrastructure-Researcher,Technical-Research,AI-Systems-Performance-Specialist,AI-Systems-Analyst,AI-Research-Specialist,AI-ML-Research-Scientist,Research-Analyst
Apply on original site ↗
Employment Type: Full-Time Work Setting: Remote Work Location: Korea Work Hours: Office hours Find out more here: About SemiAnalysis SemiAnalysis is an independent research and analysis firm specializing in the Semiconductor and AI industries. Our in-depth coverage spans the entire supply chain, from semiconductor fabrication processes to cutting-edge AI Models, software, and infrastructure. We are recognized as the leading authority on the semiconductor supply chain, with the highest concentration of industry experts within one team, and a deep-rooted passion for delving into the intricacies. We’re a global team of over 50 analysts, each with extensive networks across the semiconductor supply chain and AI ecosystem, publishing industry‑shaping articles while participating in 40+ conferences annually. Our newsletter reaches more than 200 000 subscribers worldwide, including senior management and C‑suite leaders at the leading semiconductor and AI companies. We also offer three core products: - Industry Models – we develop and publish industry models on accelerator shipments, datacenter demand and supply, GPU total cost of ownership, and more. We work with hyperscalers, neoclouds, many of the world’s largest hedge funds, and government agencies. - Core Research – our public equity markets product, geared towards financial investors, distills our deep technical research and knowledge into key insights on technology and product trends. - Consulting and Technical Due Diligence – We conduct custom research and project work to guide key strategic and investment decisions for the largest private‑equity funds, leading venture‑capital firms, companies across the AI ecosystem, and government agencies. 1) Position Overview We are seeking an AI System Performance Analyst to model the inference and training performance of AI accelerators and rack-scale systems across real-world model workloads. This role sits at the intersection of computer architecture, machine-learning systems, and market analysis. Your work will directly support the development of our Inference Simulator , InferenceX , Tokenomics Model , and AI Cloud TCO research. The central question you will answer repeatedly and rigorously is: For a given model, context length, latency target, and parallelism strategy, how many tokens per second can each accelerator and system actually deliver—and what does each token ultimately cost? You will translate chip- and system-level performance characteristics into defensible technical and economic conclusions for institutional investors, hyperscalers, semiconductor companies, and other industry participants. This role is location-flexible. Candidates based in Seoul or the broader APAC region are preferred, but location is not a requirement. 2) Responsibilities Inference Performance Modeling - Build and extend first-principles performance models for large language model inference. - Model the differences between prefill and decode workloads. - Analyze arithmetic intensity, compute utilization, memory traffic, and roofline performance limits. - Model KV-cache capacity, memory-bandwidth requirements, and context-length scaling. - Evaluate batching behavior and the trade-offs between throughput, latency, and interactivity. - Develop performance curves across different service-level objectives and deployment configurations. Accelerator and Hardware Analysis - Model performance across NVIDIA and AMD GPUs, Google TPUs, AWS Trainium, and emerging AI accelerators. - Compare accelerator architectures based on compute throughput, memory bandwidth, memory capacity, on-chip SRAM, interconnect, and system topology. - Assess how hardware design choices affect real-world inference and training performance. - Evaluate scale-up and scale-out limitations across chips, nodes, racks, and datacenter clusters. - Benchmark the relative strengths and weaknesses of heterogeneous accelerator platforms. Model Ar

← All remote jobs