References
Austin, Jacob, Sholto Douglas, Roy Frostig, et al. 2025. How to
Scale Your Model: A Systems View of LLMs on TPUs. Online book. https://jax-ml.github.io/scaling-book/.
Ba, Jimmy Lei, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016. Layer
Normalization. https://doi.org/10.48550/arXiv.1607.06450.
Baik, Jinho, and Jack W. Silverstein. 2006. “Eigenvalues of Large
Sample Covariance Matrices of Spiked Population Models.”
Journal of Multivariate Analysis 97 (6): 1382–408. https://doi.org/10.1016/j.jmva.2005.08.003.
Baraniuk, Richard, Mark Davenport, Ronald DeVore, and Michael Wakin.
2008. “A Simple Proof of the Restricted Isometry Property for
Random Matrices.” Constructive Approximation 28 (3):
253–63. https://doi.org/10.1007/s00365-007-9003-x.
Baur, Walter, and Volker Strassen. 1983. “The Complexity of
Partial Derivatives.” Theoretical Computer Science 22
(3): 317–30. https://doi.org/10.1016/0304-3975(83)90110-X.
Baydin, Atilim Gunes, Barak A. Pearlmutter, Alexey Andreyevich Radul,
and Jeffrey Mark Siskind. 2018. “Automatic Differentiation in
Machine Learning: A Survey.” Journal of Machine Learning
Research 18 (153): 1–43. https://www.jmlr.org/papers/v18/17-468.html.
Bernstein, Jeremy, and Laker Newhouse. 2024. Old Optimizer, New
Norm: An Anthology. https://doi.org/10.48550/arXiv.2409.20325.
Bottou, Léon, Frank E. Curtis, and Jorge Nocedal. 2018.
“Optimization Methods for Large-Scale Machine Learning.”
SIAM Review 60 (2): 223–311. https://doi.org/10.1137/16M1080173.
Chan, Tony F., Gene H. Golub, and Randall J. LeVeque. 1983.
“Algorithms for Computing the Sample Variance: Analysis and
Recommendations.” The American Statistician 37 (3):
242–47. https://doi.org/10.1080/00031305.1983.10483115.
Chen, Tianqi, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016.
Training Deep Nets with Sublinear Memory Cost. https://doi.org/10.48550/arXiv.1604.06174.
Chernoff, Herman. 1952. “A Measure of Asymptotic Efficiency for
Tests of a Hypothesis Based on the Sum of Observations.” The
Annals of Mathematical Statistics 23 (4): 493–507. https://doi.org/10.1214/aoms/1177729330.
Chizat, Lénaïc, Edouard Oyallon, and Francis Bach. 2019. “On Lazy
Training in Differentiable Programming.” Advances in Neural
Information Processing Systems 32. https://papers.nips.cc/paper/2019/hash/ae614c557843b1df326cb29c57225459-Abstract.html.
Cho, Youngmin, and Lawrence K. Saul. 2009. “Kernel Methods for
Deep Learning.” Advances in Neural Information Processing
Systems 22, 342–50. https://proceedings.neurips.cc/paper/2009/hash/5751ec3e9a4feab575962e78e006250d-Abstract.html.
Cohen, Jeremy M., Simran Kaur, Yuanzhi Li, J. Zico Kolter, and Ameet
Talwalkar. 2021. “Gradient Descent on Neural Networks Typically
Occurs at the Edge of Stability.” International Conference on
Learning Representations. https://openreview.net/forum?id=jh-rTtvkGeM.
Connolly, Michael P., Nicholas J. Higham, and Theo Mary. 2021.
“Stochastic Rounding and Its Probabilistic Backward Error
Analysis.” SIAM Journal on Scientific Computing 43 (1):
A566–85. https://doi.org/10.1137/20M1334796.
Dao, Tri, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.
2022. “FlashAttention: Fast and Memory-Efficient
Exact Attention with IO-Awareness.” Advances in
Neural Information Processing Systems 35, 16344–59. https://proceedings.neurips.cc/paper_files/paper/2022/hash/67d57c32e20fd0a7a302cb81d36e40d5-Abstract-Conference.html.
Glorot, Xavier, and Yoshua Bengio. 2010. “Understanding the
Difficulty of Training Deep Feedforward Neural Networks.”
Proceedings of the Thirteenth International Conference on Artificial
Intelligence and Statistics, 249–56. https://proceedings.mlr.press/v9/glorot10a.html.
Goodfellow, Ian, Yoshua Bengio, and Aaron Courville. 2016. Deep
Learning. MIT Press. https://www.deeplearningbook.org/.
Griewank, Andreas. 1992. “Achieving Logarithmic Growth of Temporal
and Spatial Complexity in Reverse Automatic Differentiation.”
Optimization Methods and Software 1 (1): 35–54. https://doi.org/10.1080/10556789208805505.
Griewank, Andreas, and Andrea Walther. 2008. Evaluating Derivatives:
Principles and Techniques of Algorithmic Differentiation. 2nd ed.
Society for Industrial; Applied Mathematics. https://doi.org/10.1137/1.9780898717761.
Gupta, Vineet, Tomer Koren, and Yoram Singer. 2018. “Shampoo:
Preconditioned Stochastic Tensor Optimization.” Proceedings
of the 35th International Conference on Machine Learning,
Proceedings of machine learning research, vol. 80: 1842–50. https://proceedings.mlr.press/v80/gupta18a.html.
He, Kaiming, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015.
“Delving Deep into Rectifiers: Surpassing Human-Level Performance
on ImageNet Classification.” Proceedings of the IEEE
International Conference on Computer Vision. https://doi.org/10.1109/ICCV.2015.123.
Hennig, Philipp. 2015. “Probabilistic Interpretation of Linear
Solvers.” SIAM Journal on Optimization 25 (1): 234–60.
https://doi.org/10.1137/140955501.
Hennig, Philipp, and Martin Kiefel. 2013. “Quasi-Newton Methods: A
New Direction.” Journal of Machine Learning Research 14:
843–65. https://www.jmlr.org/papers/v14/hennig13a.html.
Hennig, Philipp, Marvin Pförtner, and Tim Weiland. 2026.
Probabilistic Numerics: Computation Is Machine Learning.
Tutorial at the 43rd International Conference on Machine Learning. https://icml.cc/Downloads/2026.
Higham, Nicholas J. 2002. Accuracy and Stability of Numerical
Algorithms. 2nd ed. Society for Industrial; Applied Mathematics. https://doi.org/10.1137/1.9780898718027.
Hoeffding, Wassily. 1963. “Probability Inequalities for Sums of
Bounded Random Variables.” Journal of the American
Statistical Association 58 (301): 13–30. https://doi.org/10.1080/01621459.1963.10500830.
IEEE Standard for Floating-Point Arithmetic,
Pub. L. Nos. IEEE Std 754-2019 (2019). https://doi.org/10.1109/IEEESTD.2019.8766229.
Ioffe, Sergey, and Christian Szegedy. 2015. “Batch Normalization:
Accelerating Deep Network Training by Reducing Internal Covariate
Shift.” Proceedings of the 32nd International Conference on
Machine Learning, Proceedings of machine learning research, vol.
37: 448–56. https://proceedings.mlr.press/v37/ioffe15.html.
Jacot, Arthur, Franck Gabriel, and Clément Hongler. 2018. “Neural
Tangent Kernel: Convergence and Generalization in Neural
Networks.” Advances in Neural Information Processing Systems
31. https://papers.nips.cc/paper/2018/hash/5a4be1fa34e62bb8a6ec6b7ae4572128-Abstract.html.
Jordan, Keller. 2024. Muon: An Optimizer for Hidden Layers in Neural
Networks. https://kellerjordan.github.io/posts/muon/.
Kingma, Diederik P., and Jimmy Ba. 2015. “Adam: A Method for
Stochastic Optimization.” International Conference on
Learning Representations. https://doi.org/10.48550/arXiv.1412.6980.
Koloskova, Anastasia, Hadrien Hendrikx, and Sebastian U. Stich. 2023.
“Revisiting Gradient Clipping: Stochastic Bias and Tight
Convergence Guarantees.” Proceedings of the 40th
International Conference on Machine Learning, Proceedings of
machine learning research, vol. 202: 17343–63. https://proceedings.mlr.press/v202/koloskova23a.html.
Koltchinskii, Vladimir, and Karim Lounici. 2017. “Concentration
Inequalities and Moment Bounds for Sample Covariance Operators.”
Bernoulli 23 (1): 110–33. https://doi.org/10.3150/15-BEJ730.
Kunstner, Frederik, Lukas Balles, and Philipp Hennig. 2019.
“Limitations of the Empirical Fisher Approximation for Natural
Gradient Descent.” Advances in Neural Information Processing
Systems 32. https://proceedings.neurips.cc/paper/2019/hash/46a558d97954d0692411c861cf78ef79-Abstract.html.
Lee, Jaehoon, Lechao Xiao, Samuel S. Schoenholz, et al. 2019.
“Wide Neural Networks of Any Depth Evolve as Linear Models Under
Gradient Descent.” Advances in Neural Information Processing
Systems 32. https://papers.nips.cc/paper/2019/hash/0d1a9651497a38d8b1c3871c84528bd4-Abstract.html.
Linnainmaa, Seppo. 1970. “The Representation of the Cumulative
Rounding Error of an Algorithm as a Taylor Expansion of the Local
Rounding Errors.” Master’s thesis, University of Helsinki. https://kansalliskirjasto.finna.fi/Record/helka.9933382303506253.
Liu, Jingyuan et al. 2025. Muon Is
Scalable for LLM Training. https://doi.org/10.48550/arXiv.2502.16982.
Loshchilov, Ilya, and Frank Hutter. 2019. “Decoupled Weight Decay
Regularization.” International Conference on Learning
Representations. https://doi.org/10.48550/arXiv.1711.05101.
Marchenko, V. A., and L. A. Pastur. 1967. “Distribution of
Eigenvalues for Some Sets of Random Matrices.” Mathematics of
the USSR-Sbornik 1 (4): 457–83. https://doi.org/10.1070/SM1967v001n04ABEH001994.
Martens, James. 2020. “New Insights and Perspectives on the
Natural Gradient Method.” Journal of Machine Learning
Research 21 (146): 1–76. https://www.jmlr.org/papers/v21/17-678.html.
Micikevicius, Paulius, Dusan Stosic, Neil Burgess, et al. 2022.
“FP8 Formats for Deep Learning.” arXiv
Preprint arXiv:2209.05433, ahead of print. https://doi.org/10.48550/arXiv.2209.05433.
Miyato, Takeru, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida.
2018. “Spectral Normalization for Generative Adversarial
Networks.” International Conference on Learning
Representations. https://openreview.net/forum?id=B1QRgziT-.
Nesterov, Yurii. 2004. Introductory Lectures on Convex Optimization:
A Basic Course. Vol. 87. Applied Optimization. Kluwer Academic
Publishers. https://doi.org/10.1007/978-1-4419-8853-9.
Pearlmutter, Barak A. 1994. “Fast Exact Multiplication by the
Hessian.” Neural Computation 6 (1): 147–60. https://doi.org/10.1162/neco.1994.6.1.147.
Pennington, Jeffrey, Samuel Schoenholz, and Surya Ganguli. 2017.
“Resurrecting the Sigmoid in Deep Learning Through Dynamical
Isometry: Theory and Practice.” Advances in Neural
Information Processing Systems 30. https://papers.nips.cc/paper/2017/hash/d9fc0cdb67638d50f411432d0d41d0ba-Abstract.html.
Polyak, Boris T. 1964. “Some Methods of Speeding up the
Convergence of Iteration Methods.” USSR Computational
Mathematics and Mathematical Physics 4 (5): 1–17. https://doi.org/10.1016/0041-5553(64)90137-5.
Poole, Ben, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and
Surya Ganguli. 2016. “Exponential Expressivity in Deep Neural
Networks Through Transient Chaos.” Advances in Neural
Information Processing Systems 29. https://papers.nips.cc/paper/2016/hash/148510031349642de5ca0c544f31b2ef-Abstract.html.
Robbins, Herbert, and Sutton Monro. 1951. “A Stochastic
Approximation Method.” The Annals of Mathematical
Statistics 22 (3): 400–407. https://doi.org/10.1214/aoms/1177729586.
Roberts, Daniel A., and Sho Yaida. 2022. The Principles of Deep
Learning Theory: An Effective Theory Approach to Understanding Neural
Networks. Cambridge University Press. https://doi.org/10.1017/9781009023405.
Santurkar, Shibani, Dimitris Tsipras, Andrew Ilyas, and Aleksander
Madry. 2018. “How Does Batch Normalization Help
Optimization?” Advances in Neural Information Processing
Systems 31. https://papers.neurips.cc/paper/2018/hash/905056c1ac1dad141560467e0a99e1cf-Abstract.html.
Saxe, Andrew M., James L. McClelland, and Surya Ganguli. 2014.
“Exact Solutions to the Nonlinear Dynamics of Learning in Deep
Linear Neural Networks.” International Conference on Learning
Representations. https://doi.org/10.48550/arXiv.1312.6120.
Schmidt, Mark. 2026. Is Numerical Optimization Theory Irrelevant to
Machine Learning Practice in 2026? Tutorial at the 43rd
International Conference on Machine Learning. https://www.cs.ubc.ca/~schmidtm/Documents/2026_ICML_Tutorial.pdf.
Schoenholz, Samuel S., Justin Gilmer, Surya Ganguli, and Jascha
Sohl-Dickstein. 2017. “Deep Information Propagation.”
International Conference on Learning Representations. https://openreview.net/forum?id=H1W1UN9gg.
Simsekli, Umut, Levent Sagun, and Mert Gurbuzbalaban. 2019. “A
Tail-Index Analysis of Stochastic Gradient Noise in Deep Neural
Networks.” Proceedings of the 36th International Conference
on Machine Learning, Proceedings of machine learning research, vol.
97: 5827–37. https://proceedings.mlr.press/v97/simsekli19a.html.
Tazi, Nouamane, Ferdinand Mom, Haojun Zhao, et al. 2025. The
Ultra-Scale Playbook: Training LLMs on GPU Clusters. Online book.
https://huggingface.co/spaces/nanotron/ultrascale-playbook.
Vaswani, Ashish, Noam Shazeer, Niki Parmar, et al. 2017.
“Attention Is All You Need.” Advances in Neural
Information Processing Systems 30. https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html.
Vershynin, Roman. 2026. High-Dimensional Probability: An
Introduction with Applications in Data Science. 2nd ed. Cambridge
Series in Statistical and Probabilistic Mathematics. Cambridge
University Press. https://www.math.uci.edu/~rvershyn/papers/HDP-book/HDP-book.html.
Vyas, Nikhil, Depen Morwani, Rosie Zhao, et al. 2024.
SOAP: Improving and Stabilizing Shampoo
Using Adam. https://doi.org/10.48550/arXiv.2409.11321.
Welford, B. P. 1962. “Note on a Method for Calculating Corrected
Sums of Squares and Products.” Technometrics 4 (3):
419–20. https://doi.org/10.1080/00401706.1962.10490022.
Williams, Christopher K. I. 1997. “Computing with Infinite
Networks.” Advances in Neural Information Processing Systems
9, 295–301. https://papers.nips.cc/paper/1996/hash/ae5e3ce40e0404a45ecacaaf05e5f735-Abstract.html.
Williams, Samuel, Andrew Waterman, and David Patterson. 2009.
Roofline: An Insightful Visual Performance Model for Multicore
Architectures. UCB/EECS-2008-134. University of California,
Berkeley. https://digicoll.lib.berkeley.edu/record/136692.
Zandieh, Amir, Majid Daliri, Majid Hadian, and Vahab Mirrokni. 2025.
TurboQuant: Online Vector Quantization with
Near-Optimal Distortion Rate. https://doi.org/10.48550/arXiv.2504.19874.
Zandieh, Amir, Majid Daliri, and Insu Han. 2024. QJL:
1-Bit Quantized JL Transform for KV Cache
Quantization with Zero Overhead. https://doi.org/10.48550/arXiv.2406.03482.
Zhang, Biao, and Rico Sennrich. 2019. “Root Mean Square Layer
Normalization.” Advances in Neural Information Processing
Systems 32. https://papers.nips.cc/paper/2019/hash/1e8a19426224ca89e83cef47f1e7f53b-Abstract.html.
Zhang, Chiyuan, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol
Vinyals. 2017. “Understanding Deep Learning Requires Rethinking
Generalization.” International Conference on Learning
Representations. https://openreview.net/forum?id=Sy8gdB9xx.