This appendix is the sanctioned bridge between the book’s mechanism-first language and the names used in papers, libraries, and search queries. A name identifies a family; it does not remove the need to state axes, precision, state, or assumptions.
| Coordinate-wise or token-wise normalization |
LayerNorm; RMSNorm |
Axes; centering; scale; \(\varepsilon\) placement; Ba et al. (2016); Zhang and Sennrich (2019) |
Section 15.1 |
| Batch-and-spatial coordinate normalization |
BatchNorm |
Running and batch statistics are distinct estimators; Ioffe and Szegedy (2015) |
Section 15.2 |
| Variance-preserving initialization |
Xavier/Glorot; He/Kaiming |
Variance convention; nonlinearity assumptions; Glorot and Bengio (2010); He et al. (2015) |
Section 14.1 |
| Diagonal moment preconditioning |
Adam; AdamW |
Moments; bias correction; decay; \(\varepsilon\) placement; Kingma and Ba (2015); Loshchilov and Hutter (2019) |
C13/C17 |
| Gradient-dominated or noise-dominated regime |
signal-to-noise regime; critical batch regime |
\(\chi_k=\sigma_k^2/\|\nabla f(w_k)\|^2\); conditional estimator, target, and finite-moment contract |
Section 13.2 |
| Similarity-weighted aggregation |
attention; self-attention |
Scale; normalized axis; mask; materialization; Vaswani et al. (2017) |
Beyond this volume |
| IO-aware exact all-pairs aggregation |
FlashAttention |
Exact map; finite-precision order; Dao et al. (2022) |
Beyond this volume |
| Matrix-aware or orthogonalized update |
Muon; Shampoo; SOAP |
Routing; iteration count; scale; state; Jordan (2024); Gupta et al. (2018); Vyas et al. (2024) |
Section 17.1 |
| Quantized random projection or update sketch |
QJL; TurboQuant |
Quantizer; estimator; sketch size; hardware endpoint; Zandieh et al. (2024); Zandieh et al. (2025) |
Section 17.5 |
| Reverse accumulation on a computation graph |
reverse-mode automatic differentiation; backpropagation |
JVP/VJP orientation; fan-out sums; retained primal state |
C11 |
The unit column records where the full entry lives, not where the literature name is owned. A specific published artifact carries its primary citation in the row; a generic technique may instead point to the unit that owns its contract. Search aliases and those citations are tested in CI.
Ba, Jimmy Lei, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016.
Layer Normalization.
https://doi.org/10.48550/arXiv.1607.06450.
Dao, Tri, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022.
“FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.” Advances in Neural Information Processing Systems 35, 16344–59.
https://proceedings.neurips.cc/paper_files/paper/2022/hash/67d57c32e20fd0a7a302cb81d36e40d5-Abstract-Conference.html.
Glorot, Xavier, and Yoshua Bengio. 2010.
“Understanding the Difficulty of Training Deep Feedforward Neural Networks.” Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, 249–56.
https://proceedings.mlr.press/v9/glorot10a.html.
Gupta, Vineet, Tomer Koren, and Yoram Singer. 2018.
“Shampoo: Preconditioned Stochastic Tensor Optimization.” Proceedings of the 35th International Conference on Machine Learning, Proceedings of machine learning research, vol. 80: 1842–50.
https://proceedings.mlr.press/v80/gupta18a.html.
He, Kaiming, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015.
“Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification.” Proceedings of the IEEE International Conference on Computer Vision.
https://doi.org/10.1109/ICCV.2015.123.
Ioffe, Sergey, and Christian Szegedy. 2015.
“Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift.” Proceedings of the 32nd International Conference on Machine Learning, Proceedings of machine learning research, vol. 37: 448–56.
https://proceedings.mlr.press/v37/ioffe15.html.
Jordan, Keller. 2024.
Muon: An Optimizer for Hidden Layers in Neural Networks.
https://kellerjordan.github.io/posts/muon/.
Kingma, Diederik P., and Jimmy Ba. 2015.
“Adam: A Method for Stochastic Optimization.” International Conference on Learning Representations.
https://doi.org/10.48550/arXiv.1412.6980.
Loshchilov, Ilya, and Frank Hutter. 2019.
“Decoupled Weight Decay Regularization.” International Conference on Learning Representations.
https://doi.org/10.48550/arXiv.1711.05101.
Vaswani, Ashish, Noam Shazeer, Niki Parmar, et al. 2017.
“Attention Is All You Need.” Advances in Neural Information Processing Systems 30.
https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html.
Vyas, Nikhil, Depen Morwani, Rosie Zhao, et al. 2024.
SOAP: Improving and Stabilizing Shampoo Using Adam.
https://doi.org/10.48550/arXiv.2409.11321.
Zandieh, Amir, Majid Daliri, Majid Hadian, and Vahab Mirrokni. 2025.
TurboQuant: Online Vector Quantization with Near-Optimal Distortion Rate.
https://doi.org/10.48550/arXiv.2504.19874.
Zandieh, Amir, Majid Daliri, and Insu Han. 2024.
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead.
https://doi.org/10.48550/arXiv.2406.03482.
Zhang, Biao, and Rico Sennrich. 2019.
“Root Mean Square Layer Normalization.” Advances in Neural Information Processing Systems 32.
https://papers.nips.cc/paper/2019/hash/1e8a19426224ca89e83cef47f1e7f53b-Abstract.html.