Course schedule
The titles name the broader topic in each lecture, and the questions identify the issue that organizes it. The 27 lectures run from September 1 through December 8. Fall Reading Week is November 9–13, with no classes. Readings and precise theorem selections will be added later.
| Week | Lecture A | Lecture B | Coursework milestone |
|---|---|---|---|
| 1 | Tue, Sep 1 · Lecture 1 Decoder transformers and autoregressive inference How does a language model turn a prefix into a next-token distribution and generate a sequence? |
Thu, Sep 3 · Lecture 2 Learning, information, and structural generalization Why use learning, and what makes behavior on unseen inputs possible? |
Homework 0 readiness exercise; submit on Canvas by Sun, Sep 6, 11:59 p.m.; Lecture 1 model note: PDF · LaTeX |
| 2 | Tue, Sep 8 · Lecture 3 Probabilistic models, log loss, and compression Why are maximum likelihood, KL divergence, and compression different views of the same objective? |
Thu, Sep 10 · Lecture 4 Autoregressive prediction of temporal sequences Why is autoregression such a powerful way to learn sequences? |
Lecture 3 exercises: PDF (compiled 2026-09-09 21:54 UTC-06:00) · LaTeX; Homework 1 PDF · source bundle |
| 3 | Tue, Sep 15 · Lecture 5 Next-token prediction and reusable computation Why can next-token prediction recover reusable computation, and when can it fail? |
Thu, Sep 17 · Lecture 6 Learning from dependent data Is independence necessary for learning, and why is dependence also the source of predictability? |
Homework 1 due Fri, Sep 18, 11:59 p.m.; submit on Canvas; Homework 2 released |
| 4 | Tue, Sep 22 · Lecture 7 Data duplication, mixture, and misspecification When do duplication, data mixture, and curation change what is learned? |
Thu, Sep 24 · Lecture 8 Interpolation and benign overfitting Does fitting the training data, including noise, force poor generalization? |
Homework 2 in progress; research-question sketch encouraged |
| 5 | Tue, Sep 29 · Lecture 9 Implicit bias and solution selection When many predictors fit, which one does training select? |
Thu, Oct 1 · Lecture 10 Global convergence from local gradients When is local gradient information sufficient for global learning? |
Homework 2 due; Homework 3 released |
| 6 | Tue, Oct 6 · Lecture 11 Accelerated and adaptive first-order methods How can first-order descent be accelerated and better scaled? |
Thu, Oct 8 · Lecture 12 Representation learning How can learning discover features that are useful across different tasks? |
Research-question first draft due; Homework 3 in progress |
| 7 | Tue, Oct 13 · Lecture 13 Transfer from learned representations When does a learned representation help a new task? |
Thu, Oct 15 · Lecture 14 Neural networks as computational models Can neural networks represent algorithms and symbolic computation? |
Homework 3 due; Homework 4 released |
| 8 | Tue, Oct 20 · Lecture 15 Approximation by neural networks What functions can neural networks approximate, and when is approximation economical? |
Thu, Oct 22 · Lecture 16 Harnesses, state, and repeated computation Do harnesses increase computational power? |
Homework 4 in progress |
| 9 | Tue, Oct 27 · Lecture 17 Verification-guided search When does verification turn generation into efficient search? |
Thu, Oct 29 · Lecture 18 Offline imitation and log loss Can log-loss behavior cloning avoid compounding error? |
Homework 4 due; Homework 5 released |
| 10 | Tue, Nov 3 · Lecture 19 Interactive learning When does interaction improve imitation? |
Thu, Nov 5 · Lecture 20 Preference models and objectives What do pairwise preferences identify, and what are preference objectives optimizing? |
Homework 5 in progress |
| — | Tue, Nov 10 Fall Reading Week No class. |
Thu, Nov 12 Fall Reading Week No class. |
No classes November 9–13 |
| 11 | Tue, Nov 17 · Lecture 21 Optimizing learned feedback When does optimizing estimated feedback improve behavior instead of exploiting a proxy? |
Thu, Nov 19 · Lecture 22 Catastrophic forgetting and interference Why can a learner forget even when all tasks have a common solution? |
Homework 5 due; Homework 6 released |
| 12 | Tue, Nov 24 · Lecture 23 Memory in continual learning When can the past be compressed into fixed memory without losing what future learning needs? |
Thu, Nov 26 · Lecture 24 Statistical and exact learning When does statistical success mean that the underlying rule has been learned? |
Revised research-question note due; Homework 6 in progress |
| 13 | Tue, Dec 1 · Lecture 25 Denoising and score estimation Why does denoising reveal a distribution? |
Thu, Dec 3 · Lecture 26 Reverse diffusion and generative dynamics Why can reverse dynamics generate, and how do errors accumulate? |
Homework 6 due |
| Final meeting | Tue, Dec 8 · Lecture 27 Transformer training on accelerators Why does transformer training map so well to modern accelerators? |
— | New mathematical material ends |
Topics also potentially include evaluation, uncertainty, contamination, distribution shift, empirical scaling laws, curricula, and synthetic data and these are incorporated where their mathematics arises.