Skip to the content.

Deep Learning Notes

From-scratch PyTorch implementations and plain-language explanations of the core building blocks of modern deep learning. Each note pairs the intuition, the math, and a working, commented implementation you can run.

Go Board — drill rotation (priority order)

Work top-down: ~30 min read, ~15 min blank-file drill, repeat. Order reflects likelihood and value; ✅ = a clean ~30-min from-scratch drill, 💬 = understand & explain (too big or conceptual to code cold). Syntax quick-reference for all of these: PyTorch Reference Card § C.

# Subproblem Read Fit
1 Stable softmax + multi-head attention (the anchor) Transformer — Steps 3 & 5 ✅
2 Causal masking / causal self-attention Decoder — §2 ✅
3 GAT layer (attention over graph neighbours) Graph Attention — §5 ✅
4 Causal conv / TCN block Time-Series / TCN — §5 ✅
5 Cross-entropy & InfoNCE from scratch Activations/Losses — §2b ✅
6 Cross-attention (Q ≠ K/V) Decoder — §3 ✅
7 Training loop Reference Card § B · Activations/Losses §5 ✅
8 Custom Dataset / DataLoader Reference Card § A ✅
9 Encoder block (MHA + FFN + LN + residual) Transformer — Step 7 ✅*
10 ViT patch-embedding front-end Vision Transformers ✅
11 Decoder AR loop · KV cache Decoder — §6–7 💬

✅* the encoder block is clean but on the larger side (a full block) — size down to a single sub-layer if the clock is tight.

Contents

Reference

Foundations

Architectures

Applied Domain — Physical AI & Robot Learning


These notes are written to be read in short sittings. Each section stands on its own.