CS349E: Efficient ML Inference at Scale

Home Course Info

Course Description

A graduate-level course focused on principles and techniques for designing and building large-scale disaggregated AI inference systems. The course explores architectural, algorithmic, system-level, and kernel implementation techniques for serving models with low latency, high throughput, and low cost on heterogeneous machines. Topics include model optimization techniques, distributed serving, AI hardware and performance modelling, kernel programming, and AI compilers.

For course details, see the Course Info page.

We thank Gimlet Labs for graciously providing the compute infrastructure for the class.

Schedule

Week Date Topic Readings
Week 1 Lecture 1 Sept 22 (Tue) Introduction (Mark)
Lecture 2 Sept 24 (Thu) Transformers (Fred)
Week 2 Lecture 3 Sept 29 (Tue) Accuracy, Quantization, and Sparsity (Priyanka)
Lecture 4 Oct 1 (Thu) Distributed Serving (Mark)
Week 3 Lecture 5 Oct 6 (Tue) Disaggregated Serving (Zain)
Lecture 6 Oct 8 (Thu) Optimizing Autoregressive Decode (Zain)
Week 4 Lecture 7 Oct 13 (Tue) Modelling Hardware Performance (Priyanka)
Lecture 8 Oct 15 (Thu) Hardware Engines (Priyanka)
Week 5 Lecture 9 Oct 20 (Tue) Project Pitches (All)
Lecture 10 Oct 22 (Thu) KV Cache Management (Fred)
Week 6 Lecture 11 Oct 27 (Tue) Kernel Development (Fred)
Lecture 12 Oct 29 (Thu) Kernel Optimization Techniques (Priyanka)
Week 7 Nov 3 (Tue) Election Day (no class)
Lecture 13 Nov 5 (Thu) Guest Lecture
Week 8 Lecture 14 Nov 10 (Tue) ML Compilers - Frontends (Fred)
Lecture 15 Nov 12 (Thu) ML Compilers - Kernel Generation (Fred)
Week 9 Lecture 16 Nov 17 (Tue) Agentic Kernel Generation (Fred)
Lecture 17 Nov 19 (Thu) Agentic Tool Use (Fred)
Nov 23-27 Thanksgiving Recess (no class)
Week 10 Lecture 18 Dec 1 (Tue) Project Presentations (All)
Lecture 19 Dec 3 (Thu) Project Presentations (All)