Course schedule

The titles name the broader topic in each lecture, and the questions identify the issue that organizes it. The 27 lectures run from September 1 through December 8. Fall Reading Week is November 9–13, with no classes. Readings and precise theorem selections will be added later.

Week Lecture A Lecture B Coursework milestone
1 Tue, Sep 1 · Lecture 1
Decoder transformers and autoregressive inference
How does a language model turn a prefix into a next-token distribution and generate a sequence?
Thu, Sep 3 · Lecture 2
Learning, information, and structural generalization
Why use learning, and what makes behavior on unseen inputs possible?
Homework 0 readiness exercise; submit on Canvas by Sun, Sep 6, 11:59 p.m.; Lecture 1 model note: PDF · LaTeX
2 Tue, Sep 8 · Lecture 3
Probabilistic models, log loss, and compression
Why are maximum likelihood, KL divergence, and compression different views of the same objective?
Thu, Sep 10 · Lecture 4
Autoregressive prediction of temporal sequences
Why is autoregression such a powerful way to learn sequences?
Lecture 3 exercises: PDF (compiled 2026-09-09 21:54 UTC-06:00) · LaTeX; Homework 1 PDF · source bundle
3 Tue, Sep 15 · Lecture 5
Next-token prediction and reusable computation
Why can next-token prediction recover reusable computation, and when can it fail?
Thu, Sep 17 · Lecture 6
Learning from dependent data
Is independence necessary for learning, and why is dependence also the source of predictability?
Homework 1 due Fri, Sep 18, 11:59 p.m.; submit on Canvas; Homework 2 released
4 Tue, Sep 22 · Lecture 7
Data duplication, mixture, and misspecification
When do duplication, data mixture, and curation change what is learned?
Thu, Sep 24 · Lecture 8
Interpolation and benign overfitting
Does fitting the training data, including noise, force poor generalization?
Homework 2 in progress; research-question sketch encouraged
5 Tue, Sep 29 · Lecture 9
Implicit bias and solution selection
When many predictors fit, which one does training select?
Thu, Oct 1 · Lecture 10
Global convergence from local gradients
When is local gradient information sufficient for global learning?
Homework 2 due; Homework 3 released
6 Tue, Oct 6 · Lecture 11
Accelerated and adaptive first-order methods
How can first-order descent be accelerated and better scaled?
Thu, Oct 8 · Lecture 12
Representation learning
How can learning discover features that are useful across different tasks?
Research-question first draft due; Homework 3 in progress
7 Tue, Oct 13 · Lecture 13
Transfer from learned representations
When does a learned representation help a new task?
Thu, Oct 15 · Lecture 14
Neural networks as computational models
Can neural networks represent algorithms and symbolic computation?
Homework 3 due; Homework 4 released
8 Tue, Oct 20 · Lecture 15
Approximation by neural networks
What functions can neural networks approximate, and when is approximation economical?
Thu, Oct 22 · Lecture 16
Harnesses, state, and repeated computation
Do harnesses increase computational power?
Homework 4 in progress
9 Tue, Oct 27 · Lecture 17
Verification-guided search
When does verification turn generation into efficient search?
Thu, Oct 29 · Lecture 18
Offline imitation and log loss
Can log-loss behavior cloning avoid compounding error?
Homework 4 due; Homework 5 released
10 Tue, Nov 3 · Lecture 19
Interactive learning
When does interaction improve imitation?
Thu, Nov 5 · Lecture 20
Preference models and objectives
What do pairwise preferences identify, and what are preference objectives optimizing?
Homework 5 in progress
Tue, Nov 10
Fall Reading Week
No class.
Thu, Nov 12
Fall Reading Week
No class.
No classes November 9–13
11 Tue, Nov 17 · Lecture 21
Optimizing learned feedback
When does optimizing estimated feedback improve behavior instead of exploiting a proxy?
Thu, Nov 19 · Lecture 22
Catastrophic forgetting and interference
Why can a learner forget even when all tasks have a common solution?
Homework 5 due; Homework 6 released
12 Tue, Nov 24 · Lecture 23
Memory in continual learning
When can the past be compressed into fixed memory without losing what future learning needs?
Thu, Nov 26 · Lecture 24
Statistical and exact learning
When does statistical success mean that the underlying rule has been learned?
Revised research-question note due; Homework 6 in progress
13 Tue, Dec 1 · Lecture 25
Denoising and score estimation
Why does denoising reveal a distribution?
Thu, Dec 3 · Lecture 26
Reverse diffusion and generative dynamics
Why can reverse dynamics generate, and how do errors accumulate?
Homework 6 due
Final meeting Tue, Dec 8 · Lecture 27
Transformer training on accelerators
Why does transformer training map so well to modern accelerators?
New mathematical material ends

Topics also potentially include evaluation, uncertainty, contamination, distribution shift, empirical scaling laws, curricula, and synthetic data and these are incorporated where their mathematics arises.


Back to top

CMPUT 654 · Mathematical Foundations of Modern AI Systems · Fall 2026