References

Austin, Jacob, Sholto Douglas, Roy Frostig, et al. 2025. How to Scale Your Model: A Systems View of LLMs on TPUs. Online book. https://jax-ml.github.io/scaling-book/.
Ba, Jimmy Lei, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016. Layer Normalization. https://doi.org/10.48550/arXiv.1607.06450.
Baik, Jinho, and Jack W. Silverstein. 2006. “Eigenvalues of Large Sample Covariance Matrices of Spiked Population Models.” Journal of Multivariate Analysis 97 (6): 1382–408. https://doi.org/10.1016/j.jmva.2005.08.003.
Baraniuk, Richard, Mark Davenport, Ronald DeVore, and Michael Wakin. 2008. “A Simple Proof of the Restricted Isometry Property for Random Matrices.” Constructive Approximation 28 (3): 253–63. https://doi.org/10.1007/s00365-007-9003-x.
Baur, Walter, and Volker Strassen. 1983. “The Complexity of Partial Derivatives.” Theoretical Computer Science 22 (3): 317–30. https://doi.org/10.1016/0304-3975(83)90110-X.
Baydin, Atilim Gunes, Barak A. Pearlmutter, Alexey Andreyevich Radul, and Jeffrey Mark Siskind. 2018. “Automatic Differentiation in Machine Learning: A Survey.” Journal of Machine Learning Research 18 (153): 1–43. https://www.jmlr.org/papers/v18/17-468.html.
Bernstein, Jeremy, and Laker Newhouse. 2024. Old Optimizer, New Norm: An Anthology. https://doi.org/10.48550/arXiv.2409.20325.
Bottou, Léon, Frank E. Curtis, and Jorge Nocedal. 2018. “Optimization Methods for Large-Scale Machine Learning.” SIAM Review 60 (2): 223–311. https://doi.org/10.1137/16M1080173.
Chan, Tony F., Gene H. Golub, and Randall J. LeVeque. 1983. “Algorithms for Computing the Sample Variance: Analysis and Recommendations.” The American Statistician 37 (3): 242–47. https://doi.org/10.1080/00031305.1983.10483115.
Chen, Tianqi, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016. Training Deep Nets with Sublinear Memory Cost. https://doi.org/10.48550/arXiv.1604.06174.
Chernoff, Herman. 1952. “A Measure of Asymptotic Efficiency for Tests of a Hypothesis Based on the Sum of Observations.” The Annals of Mathematical Statistics 23 (4): 493–507. https://doi.org/10.1214/aoms/1177729330.
Chizat, Lénaïc, Edouard Oyallon, and Francis Bach. 2019. “On Lazy Training in Differentiable Programming.” Advances in Neural Information Processing Systems 32. https://papers.nips.cc/paper/2019/hash/ae614c557843b1df326cb29c57225459-Abstract.html.
Cho, Youngmin, and Lawrence K. Saul. 2009. “Kernel Methods for Deep Learning.” Advances in Neural Information Processing Systems 22, 342–50. https://proceedings.neurips.cc/paper/2009/hash/5751ec3e9a4feab575962e78e006250d-Abstract.html.
Cohen, Jeremy M., Simran Kaur, Yuanzhi Li, J. Zico Kolter, and Ameet Talwalkar. 2021. “Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability.” International Conference on Learning Representations. https://openreview.net/forum?id=jh-rTtvkGeM.
Connolly, Michael P., Nicholas J. Higham, and Theo Mary. 2021. “Stochastic Rounding and Its Probabilistic Backward Error Analysis.” SIAM Journal on Scientific Computing 43 (1): A566–85. https://doi.org/10.1137/20M1334796.
Dao, Tri, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.” Advances in Neural Information Processing Systems 35, 16344–59. https://proceedings.neurips.cc/paper_files/paper/2022/hash/67d57c32e20fd0a7a302cb81d36e40d5-Abstract-Conference.html.
Glorot, Xavier, and Yoshua Bengio. 2010. “Understanding the Difficulty of Training Deep Feedforward Neural Networks.” Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, 249–56. https://proceedings.mlr.press/v9/glorot10a.html.
Goodfellow, Ian, Yoshua Bengio, and Aaron Courville. 2016. Deep Learning. MIT Press. https://www.deeplearningbook.org/.
Griewank, Andreas. 1992. “Achieving Logarithmic Growth of Temporal and Spatial Complexity in Reverse Automatic Differentiation.” Optimization Methods and Software 1 (1): 35–54. https://doi.org/10.1080/10556789208805505.
Griewank, Andreas, and Andrea Walther. 2008. Evaluating Derivatives: Principles and Techniques of Algorithmic Differentiation. 2nd ed. Society for Industrial; Applied Mathematics. https://doi.org/10.1137/1.9780898717761.
Gupta, Vineet, Tomer Koren, and Yoram Singer. 2018. “Shampoo: Preconditioned Stochastic Tensor Optimization.” Proceedings of the 35th International Conference on Machine Learning, Proceedings of machine learning research, vol. 80: 1842–50. https://proceedings.mlr.press/v80/gupta18a.html.
He, Kaiming, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015. “Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification.” Proceedings of the IEEE International Conference on Computer Vision. https://doi.org/10.1109/ICCV.2015.123.
Hennig, Philipp. 2015. “Probabilistic Interpretation of Linear Solvers.” SIAM Journal on Optimization 25 (1): 234–60. https://doi.org/10.1137/140955501.
Hennig, Philipp, and Martin Kiefel. 2013. “Quasi-Newton Methods: A New Direction.” Journal of Machine Learning Research 14: 843–65. https://www.jmlr.org/papers/v14/hennig13a.html.
Hennig, Philipp, Marvin Pförtner, and Tim Weiland. 2026. Probabilistic Numerics: Computation Is Machine Learning. Tutorial at the 43rd International Conference on Machine Learning. https://icml.cc/Downloads/2026.
Higham, Nicholas J. 2002. Accuracy and Stability of Numerical Algorithms. 2nd ed. Society for Industrial; Applied Mathematics. https://doi.org/10.1137/1.9780898718027.
Hoeffding, Wassily. 1963. “Probability Inequalities for Sums of Bounded Random Variables.” Journal of the American Statistical Association 58 (301): 13–30. https://doi.org/10.1080/01621459.1963.10500830.
IEEE Standard for Floating-Point Arithmetic, Pub. L. Nos. IEEE Std 754-2019 (2019). https://doi.org/10.1109/IEEESTD.2019.8766229.
Ioffe, Sergey, and Christian Szegedy. 2015. “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift.” Proceedings of the 32nd International Conference on Machine Learning, Proceedings of machine learning research, vol. 37: 448–56. https://proceedings.mlr.press/v37/ioffe15.html.
Jacot, Arthur, Franck Gabriel, and Clément Hongler. 2018. “Neural Tangent Kernel: Convergence and Generalization in Neural Networks.” Advances in Neural Information Processing Systems 31. https://papers.nips.cc/paper/2018/hash/5a4be1fa34e62bb8a6ec6b7ae4572128-Abstract.html.
Jordan, Keller. 2024. Muon: An Optimizer for Hidden Layers in Neural Networks. https://kellerjordan.github.io/posts/muon/.
Kingma, Diederik P., and Jimmy Ba. 2015. “Adam: A Method for Stochastic Optimization.” International Conference on Learning Representations. https://doi.org/10.48550/arXiv.1412.6980.
Koloskova, Anastasia, Hadrien Hendrikx, and Sebastian U. Stich. 2023. “Revisiting Gradient Clipping: Stochastic Bias and Tight Convergence Guarantees.” Proceedings of the 40th International Conference on Machine Learning, Proceedings of machine learning research, vol. 202: 17343–63. https://proceedings.mlr.press/v202/koloskova23a.html.
Koltchinskii, Vladimir, and Karim Lounici. 2017. “Concentration Inequalities and Moment Bounds for Sample Covariance Operators.” Bernoulli 23 (1): 110–33. https://doi.org/10.3150/15-BEJ730.
Kunstner, Frederik, Lukas Balles, and Philipp Hennig. 2019. “Limitations of the Empirical Fisher Approximation for Natural Gradient Descent.” Advances in Neural Information Processing Systems 32. https://proceedings.neurips.cc/paper/2019/hash/46a558d97954d0692411c861cf78ef79-Abstract.html.
Lee, Jaehoon, Lechao Xiao, Samuel S. Schoenholz, et al. 2019. “Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent.” Advances in Neural Information Processing Systems 32. https://papers.nips.cc/paper/2019/hash/0d1a9651497a38d8b1c3871c84528bd4-Abstract.html.
Linnainmaa, Seppo. 1970. “The Representation of the Cumulative Rounding Error of an Algorithm as a Taylor Expansion of the Local Rounding Errors.” Master’s thesis, University of Helsinki. https://kansalliskirjasto.finna.fi/Record/helka.9933382303506253.
Liu, Jingyuan et al. 2025. Muon Is Scalable for LLM Training. https://doi.org/10.48550/arXiv.2502.16982.
Loshchilov, Ilya, and Frank Hutter. 2019. “Decoupled Weight Decay Regularization.” International Conference on Learning Representations. https://doi.org/10.48550/arXiv.1711.05101.
Marchenko, V. A., and L. A. Pastur. 1967. “Distribution of Eigenvalues for Some Sets of Random Matrices.” Mathematics of the USSR-Sbornik 1 (4): 457–83. https://doi.org/10.1070/SM1967v001n04ABEH001994.
Martens, James. 2020. “New Insights and Perspectives on the Natural Gradient Method.” Journal of Machine Learning Research 21 (146): 1–76. https://www.jmlr.org/papers/v21/17-678.html.
Micikevicius, Paulius, Dusan Stosic, Neil Burgess, et al. 2022. FP8 Formats for Deep Learning.” arXiv Preprint arXiv:2209.05433, ahead of print. https://doi.org/10.48550/arXiv.2209.05433.
Miyato, Takeru, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. 2018. “Spectral Normalization for Generative Adversarial Networks.” International Conference on Learning Representations. https://openreview.net/forum?id=B1QRgziT-.
Nesterov, Yurii. 2004. Introductory Lectures on Convex Optimization: A Basic Course. Vol. 87. Applied Optimization. Kluwer Academic Publishers. https://doi.org/10.1007/978-1-4419-8853-9.
Pearlmutter, Barak A. 1994. “Fast Exact Multiplication by the Hessian.” Neural Computation 6 (1): 147–60. https://doi.org/10.1162/neco.1994.6.1.147.
Pennington, Jeffrey, Samuel Schoenholz, and Surya Ganguli. 2017. “Resurrecting the Sigmoid in Deep Learning Through Dynamical Isometry: Theory and Practice.” Advances in Neural Information Processing Systems 30. https://papers.nips.cc/paper/2017/hash/d9fc0cdb67638d50f411432d0d41d0ba-Abstract.html.
Polyak, Boris T. 1964. “Some Methods of Speeding up the Convergence of Iteration Methods.” USSR Computational Mathematics and Mathematical Physics 4 (5): 1–17. https://doi.org/10.1016/0041-5553(64)90137-5.
Poole, Ben, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli. 2016. “Exponential Expressivity in Deep Neural Networks Through Transient Chaos.” Advances in Neural Information Processing Systems 29. https://papers.nips.cc/paper/2016/hash/148510031349642de5ca0c544f31b2ef-Abstract.html.
Robbins, Herbert, and Sutton Monro. 1951. “A Stochastic Approximation Method.” The Annals of Mathematical Statistics 22 (3): 400–407. https://doi.org/10.1214/aoms/1177729586.
Roberts, Daniel A., and Sho Yaida. 2022. The Principles of Deep Learning Theory: An Effective Theory Approach to Understanding Neural Networks. Cambridge University Press. https://doi.org/10.1017/9781009023405.
Santurkar, Shibani, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry. 2018. “How Does Batch Normalization Help Optimization?” Advances in Neural Information Processing Systems 31. https://papers.neurips.cc/paper/2018/hash/905056c1ac1dad141560467e0a99e1cf-Abstract.html.
Saxe, Andrew M., James L. McClelland, and Surya Ganguli. 2014. “Exact Solutions to the Nonlinear Dynamics of Learning in Deep Linear Neural Networks.” International Conference on Learning Representations. https://doi.org/10.48550/arXiv.1312.6120.
Schmidt, Mark. 2026. Is Numerical Optimization Theory Irrelevant to Machine Learning Practice in 2026? Tutorial at the 43rd International Conference on Machine Learning. https://www.cs.ubc.ca/~schmidtm/Documents/2026_ICML_Tutorial.pdf.
Schoenholz, Samuel S., Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein. 2017. “Deep Information Propagation.” International Conference on Learning Representations. https://openreview.net/forum?id=H1W1UN9gg.
Simsekli, Umut, Levent Sagun, and Mert Gurbuzbalaban. 2019. “A Tail-Index Analysis of Stochastic Gradient Noise in Deep Neural Networks.” Proceedings of the 36th International Conference on Machine Learning, Proceedings of machine learning research, vol. 97: 5827–37. https://proceedings.mlr.press/v97/simsekli19a.html.
Tazi, Nouamane, Ferdinand Mom, Haojun Zhao, et al. 2025. The Ultra-Scale Playbook: Training LLMs on GPU Clusters. Online book. https://huggingface.co/spaces/nanotron/ultrascale-playbook.
Vaswani, Ashish, Noam Shazeer, Niki Parmar, et al. 2017. “Attention Is All You Need.” Advances in Neural Information Processing Systems 30. https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html.
Vershynin, Roman. 2026. High-Dimensional Probability: An Introduction with Applications in Data Science. 2nd ed. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press. https://www.math.uci.edu/~rvershyn/papers/HDP-book/HDP-book.html.
Vyas, Nikhil, Depen Morwani, Rosie Zhao, et al. 2024. SOAP: Improving and Stabilizing Shampoo Using Adam. https://doi.org/10.48550/arXiv.2409.11321.
Welford, B. P. 1962. “Note on a Method for Calculating Corrected Sums of Squares and Products.” Technometrics 4 (3): 419–20. https://doi.org/10.1080/00401706.1962.10490022.
Williams, Christopher K. I. 1997. “Computing with Infinite Networks.” Advances in Neural Information Processing Systems 9, 295–301. https://papers.nips.cc/paper/1996/hash/ae5e3ce40e0404a45ecacaaf05e5f735-Abstract.html.
Williams, Samuel, Andrew Waterman, and David Patterson. 2009. Roofline: An Insightful Visual Performance Model for Multicore Architectures. UCB/EECS-2008-134. University of California, Berkeley. https://digicoll.lib.berkeley.edu/record/136692.
Zandieh, Amir, Majid Daliri, Majid Hadian, and Vahab Mirrokni. 2025. TurboQuant: Online Vector Quantization with Near-Optimal Distortion Rate. https://doi.org/10.48550/arXiv.2504.19874.
Zandieh, Amir, Majid Daliri, and Insu Han. 2024. QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead. https://doi.org/10.48550/arXiv.2406.03482.
Zhang, Biao, and Rico Sennrich. 2019. “Root Mean Square Layer Normalization.” Advances in Neural Information Processing Systems 32. https://papers.nips.cc/paper/2019/hash/1e8a19426224ca89e83cef47f1e7f53b-Abstract.html.
Zhang, Chiyuan, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. 2017. “Understanding Deep Learning Requires Rethinking Generalization.” International Conference on Learning Representations. https://openreview.net/forum?id=Sy8gdB9xx.