Research Analyst - AI System Performance Modelling

 Posted 15 hours ago
     
2-5 years experience
Apply Now

Please mention DailyRemote when applying

AI Summary

The role involves building first-principles performance models for AI accelerators and rack-scale systems to analyze inference and training workloads. The analyst will translate technical hardware characteristics into economic conclusions and TCO research for institutional clients.

Employment Type: Full-Time
Work Setting: Remote
Work Location: Korea
Work Hours: Office hours
Find out more here: https://semianalysis.com

About SemiAnalysis

SemiAnalysis is an independent research and analysis firm specializing in the Semiconductor and AI industries. Our in-depth coverage spans the entire supply chain, from semiconductor fabrication processes to cutting-edge AI Models, software, and infrastructure. We are recognized as the leading authority on the semiconductor supply chain, with the highest concentration of industry experts within one team, and a deep-rooted passion for delving into the intricacies.

We’re a global team of over 50 analysts, each with extensive networks across the semiconductor supply chain and AI ecosystem, publishing industry‑shaping articles while participating in 40+ conferences annually.

Our newsletter reaches more than 200 000 subscribers worldwide, including senior management and C‑suite leaders at the leading semiconductor and AI companies.

We also offer three core products:

  • Industry Models – we develop and publish industry models on accelerator shipments, datacenter demand and supply, GPU total cost of ownership, and more. We work with hyperscalers, neoclouds, many of the world’s largest hedge funds, and government agencies.

  • Core Research – our public equity markets product, geared towards financial investors, distills our deep technical research and knowledge into key insights on technology and product trends.

  • Consulting and Technical Due Diligence – We conduct custom research and project work to guide key strategic and investment decisions for the largest private‑equity funds, leading venture‑capital firms, companies across the AI ecosystem, and government agencies.

1) Position Overview

We are seeking an AI System Performance Analyst to model the inference and training performance of AI accelerators and rack-scale systems across real-world model workloads.

This role sits at the intersection of computer architecture, machine-learning systems, and market analysis. Your work will directly support the development of our Inference Simulator, InferenceX, Tokenomics Model, and AI Cloud TCO research.

The central question you will answer repeatedly and rigorously is:

For a given model, context length, latency target, and parallelism strategy, how many tokens per second can each accelerator and system actually deliver—and what does each token ultimately cost?

You will translate chip- and system-level performance characteristics into defensible technical and economic conclusions for institutional investors, hyperscalers, semiconductor companies, and other industry participants.

This role is location-flexible. Candidates based in Seoul or the broader APAC region are preferred, but location is not a requirement.

2) Responsibilities

Inference Performance Modeling

  • Build and extend first-principles performance models for large language model inference.

  • Model the differences between prefill and decode workloads.

  • Analyze arithmetic intensity, compute utilization, memory traffic, and roofline performance limits.

  • Model KV-cache capacity, memory-bandwidth requirements, and context-length scaling.

  • Evaluate batching behavior and the trade-offs between throughput, latency, and interactivity.

  • Develop performance curves across different service-level objectives and deployment configurations.

Accelerator and Hardware Analysis

  • Model performance across NVIDIA and AMD GPUs, Google TPUs, AWS Trainium, and emerging AI accelerators.

  • Compare accelerator architectures based on compute throughput, memory bandwidth, memory capacity, on-chip SRAM, interconnect, and system topology.

  • Assess how hardware design choices affect real-world inference and training performance.

  • Evaluate scale-up and scale-out limitations across chips, nodes, racks, and datacenter clusters.

  • Benchmark the relative strengths and weaknesses of heterogeneous accelerator platforms.

Model Architecture and Workload Analysis

  • Model dense transformer and mixture-of-experts architectures.

  • Analyze long-context, reasoning, multimodal, and other computationally demanding workloads.

  • Evaluate speculative decoding, quantization formats, sparsity, and other inference-optimization techniques.

  • Model disaggregated prefill and decode serving architectures.

  • Assess how model architecture, parameter count, active parameter count, sequence length, and precision affect system performance.

  • Track changes in model architecture that materially influence hardware requirements and deployment economics.

Parallelism and Rack-Scale Systems

  • Analyze tensor, pipeline, expert, and data-parallel strategies.

  • Evaluate how parallelism schemes map onto different accelerator and rack-scale architectures.

  • Model collective communication overheads, synchronization costs, and scaling efficiency.

  • Assess rack-scale systems such as NVL72-class platforms and comparable architectures.

  • Evaluate the effects of scale-up fabrics, network topology, link bandwidth, and congestion on delivered performance.

  • Identify system bottlenecks that prevent theoretical accelerator performance from being achieved in production.

Benchmarking and Validation

  • Validate performance models against published and independently gathered benchmarks.

  • Analyze benchmark data from frameworks and deployments using vLLM, SGLang, TensorRT-LLM, PyTorch, JAX, and similar platforms.

  • Reconcile differences between theoretical performance, vendor claims, benchmark results, and production deployments.

  • Identify methodological weaknesses, hidden assumptions, and configuration differences across benchmark datasets.

  • Develop reproducible benchmarking and analytical workflows.

  • Continuously refine model assumptions using new hardware disclosures, software improvements, and real-world performance data.

Performance Economics and TCO

  • Translate technical performance into economic metrics.

  • Model tokens per second per accelerator, server, rack, megawatt, and dollar of capital expenditure.

  • Evaluate tokens per watt and the impact of utilization on operating costs.

  • Connect hardware performance to datacenter power, cooling, networking, and infrastructure requirements.

  • Contribute performance inputs to AI cloud TCO and token-cost models.

  • Assess the economic implications of hardware selection, model architecture, latency targets, and deployment strategy.

  • Help institutional clients understand the cost curves of AI training and inference.

Research and Collaboration

  • Publish technical deep dives on accelerator, inference, training, and system performance.

  • Serve as a technical authority during client calls, briefings, and research discussions.

  • Communicate complex performance findings clearly to both engineering and investment audiences.

  • Collaborate with SemiAnalysis’ accelerator, networking, memory, datacenter, and market analysts.

  • Connect chip-level performance analysis to system-level, financial, and industry conclusions.

  • Contribute to major newsletters, research reports, client projects, and proprietary analytical products.

3) Requirements

  • 2–5+ years of experience in ML systems engineering, accelerator or GPU performance engineering, computer architecture, or performance-focused technical analysis.

  • Strong quantitative understanding of transformer inference and training workloads.

  • Ability to calculate and model:

    • FLOPs per token.

    • Memory traffic per token.

    • KV-cache capacity and bandwidth requirements.

    • Arithmetic intensity.

    • Batching effects.

    • Context-length scaling.

    • Model-architecture and hardware interactions.

  • Working knowledge of modern accelerator architectures and memory systems.

  • Understanding of HBM bandwidth and capacity trade-offs, on-chip SRAM, memory hierarchy, and data movement.

  • Familiarity with scale-up interconnects such as NVLink, UALink, or comparable technologies.

  • Understanding of scale-out networking and distributed-system performance.

  • Proficiency in Python for performance modeling, data analysis, simulation, and reproducible analytical tooling.

  • Hands-on familiarity with at least one serving or training framework, such as:

    • vLLM.

    • SGLang.

    • TensorRT-LLM.

    • PyTorch.

    • JAX.

  • Ability to independently define a modeling problem, identify the necessary data, build the analysis, validate the findings, and produce a defensible conclusion.

  • Strong written and verbal communication skills.

  • Ability to explain highly technical concepts clearly to both engineering and investor audiences.

  • Self-driven working style and the ability to operate effectively with minimal oversight.

4) Preferred Skills

  • Experience writing or optimizing GPU kernels using CUDA, Triton, HIP, or similar programming environments.

  • Experience profiling and optimizing production inference or training deployments.

  • Familiarity with kernel-level bottlenecks, operator fusion, memory access patterns, and hardware utilization.

  • Experience with training-performance modeling, including:

    • Model FLOPs utilization.

    • Parallelism-scaling efficiency.

    • Gradient and activation checkpointing.

    • Communication overhead.

    • Pipeline bubbles.

    • Failure-recovery and checkpointing overhead.

  • Experience benchmarking across heterogeneous hardware platforms.

  • Familiarity with non-NVIDIA accelerators, including TPUs, Trainium, AMD GPUs, or emerging custom silicon.

  • Understanding of model-serving infrastructure, schedulers, orchestration, and distributed inference systems.

  • Experience analyzing power consumption, datacenter infrastructure, and cost of ownership.

  • Prior published technical writing, academic research, open-source contributions, or conference presentations related to ML systems, accelerators, or computer architecture.

  • Familiarity with cloud accelerator pricing, utilization economics, and infrastructure-capacity planning.

  • Experience translating engineering performance into financial, market, or investment implications.

5) Growth Areas

This role offers the opportunity to become a leading authority on the performance and economics of AI computing systems.

Potential growth areas include:

  • Taking broader ownership of SemiAnalysis’ Inference Simulator, InferenceX, Tokenomics Model, and AI Cloud TCO research.

  • Developing proprietary methodologies for evaluating real-world accelerator and rack-scale performance.

  • Expanding coverage from inference into large-scale training, fine-tuning, reinforcement learning, and multimodal workloads.

  • Building deeper expertise in GPU kernels, compilers, serving frameworks, distributed systems, and performance optimization.

  • Developing industry-leading analysis of emerging accelerators and heterogeneous AI infrastructure.

  • Leading independent benchmarking projects across chips, servers, racks, and cloud platforms.

  • Becoming a recognized external expert on AI accelerator performance, inference economics, and token-cost modeling.

  • Publishing major research reports and presenting findings to investors, hyperscalers, semiconductor companies, and AI infrastructure providers.

  • Working closely with leading engineers, researchers, executives, and infrastructure decision-makers across the AI ecosystem.

  • Mentoring junior analysts and helping expand the firm’s ML systems and performance-analysis capabilities.

  • Progressing into a senior analyst, technical research lead, model owner, or broader AI infrastructure research leadership position.

Similar Jobs

See all Remote Software Development jobs →

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in Research Analyst

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified