Mao Lin (林茂)

ML Systems · AI Accelerators · Performance Analysis · Infrastructure

I build ML/LLM systems, AI accelerators, and performance analysis tooling. I received my Ph.D. in EECS from UC Merced, and my Master of Software Engineering and Bachelor of Computer Science from Shandong University.

Experience

Samsung

San Jose, CA, USA · 05/2026 - Present

Senior Engineer 09/2026 - Present

Working on the full zHBM AI accelerator stack, from LLM model architecture down to microarchitecture.

Research Intern 05/2026 - 08/2026

Worked on hardware/software co-design for MoE models on Samsung's AI accelerator.

ByteDance - Research Intern

05/2023 - 11/2023

Seattle/San Jose, WA/CA, USA

Optimized PyTorch memory management for distributed LLM training, reducing memory usage by 10% to 30% on models including GPT-2 and Whisper.

Uber - Software Engineer Intern

11/2022 - 02/2023

Sunnyvale, CA, USA

Analyzed production Go services and fixed more than 50 data race issues.

PNNL - Research Intern

06/2022 - 08/2022

Richland, WA, USA

Built GPU profiling and floating-point analysis tooling that found critical overflow issues in DOE applications.

Open Source Software

AccelProf

A profiling and analysis framework for various accelerator applications.

DrGPUM

Tooling for guiding memory optimization in GPU-accelerated applications.

Selected Publications

GPU Memory Optimization

AutoUVM: Automated Prefetching Framework for LLMs under UVM Oversubscription

ICCD '26

Mao Lin, Hui Feng, Xianzhong Ding, Guilherme Cox, Qian Wang, and Hyeran Jeon

The 44th IEEE International Conference on Computer Design

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing

arXiv 2026

Mao Lin, Xi Wang, Guilherme Cox, Dong Li, and Hyeran Jeon

Preprint

Forest: Access-aware GPU UVM Management

ISCA '25

Mao Lin, Yuan Feng, Guilherme Cox, and Hyeran Jeon

The 52nd Annual International Symposium on Computer Architecture

GPU Memory Profiling

LLMPROF: Identifying Performance Bottlenecks in LLM Serving Systems with Top-Down Profiling

SC '26

Tianle Zhong, Mao Lin, Hao Wu, Keren Zhou, and Geoffrey Fox

The International Conference for High Performance Computing, Networking, Storage, and Analysis (Supercomputing)

PASTA: A Modular Program Analysis Tool Framework for Accelerators

CGO '26

Mao Lin, Hyeran Jeon, and Keren Zhou

The 23rd ACM/IEEE International Symposium on Code Generation and Optimization

DrGPUM: Guiding Memory Optimization for GPU-accelerated Applications

ASPLOS '23

Mao Lin, Keren Zhou, and Pengfei Su

The 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems

Get in Touch