Twelve layers each preserve squared norm in expectation. Their product destroys one direction to \(9.6\times10^{-13}\).
The registered witness uses independent \(24\times24\) Gaussian matrices with entry variance \(1/24\). Their product has mean squared singular value \(0.795\) and largest singular value \(3.62\). An equally deep product of orthogonal factors has every singular value equal to one up to rounding. The average-moment story cannot distinguish them.
ImportantPrediction
Two initializations preserve activation variance layer by layer. Must they propagate every gradient direction comparably? Choose the control you would inspect next: another activation moment, the smallest and largest singular values, or only the parameter norm. What would falsify your choice?
with independent mean-zero weights of variance \(\sigma_w^2/n\). Under the usual wide, exchangeable, self-averaging approximation, a typical preactivation is approximately Gaussian and the second moment follows a scalar map.
Theorem 14.1 (Wide-model second-moment recursion) Under the independence and Gaussian-limit assumptions above,
For zero-bias ReLU coordinates, \(q_{\ell+1}=\sigma_w^2q_\ell/2\); the variance-critical scale is \(\sigma_w^2=2\).
Diagnostic proof
Condition on the previous coordinates. Independence and the \(1/n\) scaling make the preactivation variance the empirical second moment times \(\sigma_w^2\), plus bias variance. The wide limit replaces that empirical moment by \(q_\ell\). Symmetry gives \(\E[\max(0,\sqrt q Z)^2]=q/2\).
The theorem identifies a boundary: scale one decays by \(2^{-\ell}\), scale two is fixed, and scale three grows by \((3/2)^\ell\). It does not say that a finite network exactly follows the recursion, nor that scale two has a well-conditioned input-output Jacobian.
14.2 The moment survived; the distinction did not
A one-input moment cannot say whether two inputs remain distinguishable. Take equal-moment inputs and write their normalized inner product as \(c_\ell\). In the same wide Gaussian model, zero-bias ReLU propagation closes on a second scalar map.
Theorem 14.2 (Wide-model two-input correlation recursion) For two equal-moment inputs under the assumptions of Theorem 14.1,
For zero bias, ReLU homogeneity cancels the weight scale from this normalized map. At the variance-critical scale \(\sigma_w^2=2\), each input’s moment stays fixed while every finite initial angle tends toward alignment. If \(c_\ell=\cos\psi_\ell\) and \(\psi_\ell\) is small, then \(\psi_{\ell+1}=\psi_\ell-\psi_\ell^2/(3\pi)+O(\psi_\ell^3)\), so \(\psi_\ell\asymp 1/\ell\).
Diagnostic derivation
The two preactivations are correlated Gaussians. Integrating the product of their positive parts over the angle between them gives Equation 14.3. Expanding \(\sin\psi-\psi\cos\psi=\psi^3/3+O(\psi^5)\) near alignment gives the stated angle recurrence. Thus local pair separation is marginal, yet finite-angle separation still decays polynomially rather than remaining fixed.
The closed form in Equation 14.3 is the degree-one arc-cosine kernel introduced by Cho and Saul (2009); Williams (1997) is the Gaussian-process covariance antecedent. Here the standard name arrives only after the normalized correlation object and its wide-model assumptions are fixed.
This is the missing contract behind the scalar critical point. The structure kept every norm and slowly lost the difference between its inputs. Changing only \(\sigma_w^2\) changes the magnitude traces but not this normalized ReLU correlation trace; three weight scales do not create three pair-geometry regimes.
The forward and backward gains should also be named separately. Let \(\chi_{\parallel}\) be the derivative of the one-input moment map at its fixed point, and let \(\chi_{\perp}\) be the infinitesimal pair-separation gain. The second factor also appears in the mean-square backward recursion. They coincide for the scale-invariant piecewise-linear family used here, but that coincidence is a property of the family, not a general law connecting forward criticality to gradient preservation.
14.3 Width is relative to depth
The wide map now needs its own trust instrument. In the finite-width critical ReLU model, let \(q_\ell=n^{-1}\sum_i(h_i^{(\ell)})^2\). Conditional on the previous layer,
Theorem 14.3 (Exact finite-width moment-fluctuation control) With independent Gaussian weights of variance \(2/n\) across layers, zero bias, and ReLU coordinates,
Therefore the variance is \(5L/n+O((L/n)^2)\) only while \(L/n\) is small. If \(L,n\to\infty\) together with \(L/n\to r\), the same control approaches \(e^{5r}-1\), not a line.
Proof
For \(Y=(Z)_+^2\), symmetry gives \(\E Y=1/2\) and \(\E Y^2=\E[Z^4\mathbf 1_{Z>0}]=3/2\), hence \(\var(Y)=5/4\). Thus \(\E R_\ell=1\) and \(\var(R_\ell)=5/n\). Independent layer weights make the multipliers independent, so \(\E[(q_L/q_0)^2]=(1+5/n)^L\). Subtracting the squared mean proves the result.
Evaluate the wide two-input correlation control.
Load its seeded finite-width check.
Evaluate the exact depth-to-width fluctuation control.
Load the seeded moment-ratio checks.
Plot both contracts and test the declared perturbative boundary.
pair correlation at depth 64: 0.992491
largest Monte Carlo relative deviation: 1.756%
Figure 14.1: Left: at critical ReLU scale, each input keeps its moment while two inputs align; finite-width means follow the wide correlation map. Right: the exact moment-fluctuation variance bends away from its 5L/n tangent as depth over width grows. The Monte Carlo points check the finite-width identity rather than replacing it.
Claims c14-pair-geometry-001, c14-depth-width-001 · seeds: SHA-256-derived 3972769481 and 3834729383 · dtype: FP64 · device: CPU · estimators: 2,048-trial pair-correlation summaries and 200,000-trial population variances · pair artifact · depth/width artifact
The final wide-model correlation is \(0.992491\); the finite-width mean differs by about \(2\times10^{-6}\). Across the nine fluctuation checks, the largest Monte Carlo deviation from the exact variance is under \(1.8\%\). The seeded points validate the implementation. The exact curve carries the claim.
NoteField note: an approximation names its small parameter
Every usable limit in this book buys solvability with a declared small quantity: a step–curvature product, a noise-to-signal ratio, a covering radius, and now a depth-to-width ratio. The claim, retained order, and order of limits travel together; the claim dies where that contract stops being small. The first audit question for a clean approximation is therefore: which quantity did it assume was small, and how large is that quantity in the run on your screen?
14.4 Average control does not survive multiplication
The scalar recursion is easiest to trust after carrying it twice. Set \(q_0=1\), zero bias, and \(\sigma_w^2=2(1+\varepsilon)\). For rectified coordinates,
At \(\varepsilon=0\), the average second moment is fixed at every depth. For a small nonzero mismatch, depth compounds the local error: the relevant quantity is \(L\log(1+\varepsilon)\), not \(\varepsilon\) alone.
Now compare the directional statement. The two matrices
each have mean squared singular value one, yet \(\matr J_2\matr J_1=\matr 0\). Two layers can therefore be mean-square critical individually while their derivative product destroys every direction. Repeating the scalar recursion to depth \(L\) controls an average under its independence model; controlling the product requires directional or spectral information at the same depth.
Load the chapter-pinned signal-propagation instruments.
Plot the ReLU moment recursion at three scales.
Recreate the registered Gaussian and orthogonal matrix products.
Insert ReLU gates from one finite-width forward pass at critical scale.
Compare all three spectra and verify both committed claims.
Figure 14.2: Left: a scalar critical point separates decay from growth. Right: average preservation does not imply directional preservation; the orthogonal linear control stays isometric, while Gaussian products spread and a finite ReLU-gated product also loses rank.
The gated product is not an independent random-matrix toy. Its diagonal gates come from one forward pass through the same finite-width compositional structure whose Jacobian is being measured. At the variance-critical scale \(2/n\), the twelve layers retain between 8 and 20 active coordinates each, but the final Jacobian has only eight singular values above \(10^{-10}\). Its largest singular value is about \(2.63\) and its smallest is numerically zero. This one realization does not define a law for nonlinear Jacobians; it shows exactly where the deep-linear isometry control stops being the network itself.
14.5 Trace control is not edge control
For an input-output Jacobian \(\matr J\in\mathbb R^{m\times n}\),
Theorem 14.4 (Mean-square preservation does not imply isometry) The quantity in Equation 14.7 is the mean squared stretch over an orthonormal input basis. It can equal one while the smallest singular value is zero and the largest is \(\sqrt n\).
Proof
The trace identity follows from the SVD. The diagonal matrix \(\operatorname{diag}(\sqrt n,0,\ldots,0)\) has mean squared singular value one and the stated edges.
Dynamical isometry asks for substantially more: the singular values of the relevant Jacobian should cluster near one (Pennington et al. 2017). Orthogonal factors solve the deep-linear control exactly. Nonlinear gates, finite width, correlations, biases, normalization, and trained weights change the product; orthogonality at initialization is an instrument, not a universal cure.
The word critical is therefore incomplete until it names a rung:
Contract
Observable
What it can certify
What it cannot exclude
one-input magnitude
\(q_\ell\)
average scale survives
two inputs become indistinguishable
pair geometry
\(c_\ell\) or \(\psi_\ell\)
separation contracts, expands, or is marginal
derivative directions are ill-conditioned
derivative geometry
singular values of \(\matr J\)
directional stretch and rank
the infinite-width model is inaccurate
approximation trust
\(L/n\) with retained order
the declared expansion is controlled
training leaves the declared scaling regime
The first three rows diagnose the compositional structure. The fourth audits the theory used to describe it. In particular, an infinite-width calculation can make training linearized by suppressing feature movement under its declared parameterization and time horizon. The feature-learning branch is mapped in Beyond This Volume; no width-only diagnosis is licensed here (Jacot et al. 2018; Lee et al. 2019; Chizat et al. 2019; Roberts and Yaida 2022).
14.6 The de-branded object first
We have derived variance-preserving initialization from a moment recursion. The Rosetta appendix carries the conventional aliases for search and ties each to its fan-in/fan-out and nonlinearity assumptions. The mathematics decides which moment a name actually covers.
WarningNamed wrong answer: every activation has one critical initialization
The moment map may produce a scale-invariant critical curve, a tuned point at a zero fixed point, a half-stable boundary, or no usable critical point in the declared family. Even when a scalar critical point exists, it controls one average statistic, not finite-width fluctuations, pair geometry, spectral edges, or numerical range.
14.7 Check yourself
At zero bias with ReLU and weight scale \(2(1+\varepsilon)\), what is the moment after \(L\) layers? At the exact scalar critical point, why can two inputs still align? For \(L/n=1/8\), which is the claim: \(5L/n\), the exact finite-width formula, or its simultaneous-limit curve? Finally, what do the mean squared singular value and lower spectral edge say for \(\operatorname{diag}(\sqrt2,0)\), and which statement governs a gradient component aligned with the second coordinate?
14.8 Okay, so —
Inherited: C11’s reverse product and C08’s two spectral edges now describe signal propagation.
Changed: “critical” is separated into magnitude, pair, derivative, and approximation-trust contracts.
Instrumented: four claims compare moment recursion, pair correlation, depth-to-width fluctuations, and complete product spectra.
Established: preserving a scalar moment neither preserves finite-angle distinctions nor controls Jacobian edges; \(L/n\) marks where the wide approximation accumulates finite-width corrections.
Unresolved: Which invariances and null directions does a coordinate-wise normalization choose when it resets scale?
14.9 Sources and further reading
The Gaussian covariance antecedent is Williams (1997), and the ReLU closed form is the degree-one arc-cosine kernel of Cho and Saul (2009). Two-input signal propagation follows Poole et al. (2016) and Schoenholz et al. (2017); dynamical-isometry diagnostics follow Pennington et al. (2017). Orthogonal initialization as a controlled dynamical system appears in Saxe et al. (2014). Variance-preserving initialization was developed for saturating and rectified compositional structures by Glorot and Bengio (2010) and He et al. (2015). For a derivational companion to the moment and pair-correlation recursions, and for a systematic leading finite-width expansion whose trust parameter is depth over width, see Roberts and Yaida (2022); begin with Chapter 5, Sections 5.1, 5.4, and 5.5.
Reading order. Begin with Cho and Saul (2009) for the closed form, continue with Schoenholz et al. (2017) for the one- and two-input maps and Pennington et al. (2017) for the Jacobian spectrum, then use Roberts and Yaida (2022) to audit the depth-to-width expansion and order of limits.
14.10 Exercises
(Pencil.) Derive Equation 14.2 for an odd saturating nonlinearity and identify its small-variance linearization. Decide whether the fixed point is stable from one side, both sides, or neither. Route: Core. Estimated time: 25 minutes. Prerequisite: C04 fixed points. Deliverable: derivation plus stability classification. Hint: differentiate the moment map at the fixed point.
(Code.) Repeat the product experiment across 100 seeds. Plot distributions of trace, lower edge, and upper edge separately. Route: Core. Estimated time: 45 minutes. Prerequisite: C08 spectrum summaries. Deliverable: three matched-seed distributions and a one-paragraph verdict. Hint: do not average the lower and upper edges together.
(Pencil.) For \(\matr J=\matr D_L\matr W_L\cdots\matr D_1\matr W_1\), show that \(\operatorname{rank}(\matr J)\) cannot exceed the smallest gate rank. Explain why this bound can be loose. Route: Core. Estimated time: 20 minutes. Prerequisite: rank inequalities. Deliverable: proof and one strict example. Hint: apply the rank inequality one factor at a time.
(Audit.) Build the implication graph among moment preservation, infinitesimal pair preservation, finite-angle preservation, mean-square Jacobian preservation, dynamical isometry, and trainability. Break every invalid arrow with a counterexample or a missing assumption. In particular, locate the invalid transfer from \(\chi_\perp=1\) to finite-angle preservation and state the needed smoothness or uniformity condition. Route: Extension. Estimated time: 50 minutes. Prerequisite: this chapter. Deliverable: annotated implication graph. Hint:Equation 14.3 is marginal locally but not constant at finite angle.
(Code.) Reproduce Figure 14.1 for a smooth saturating activation. Estimate \(\chi_\parallel\) and \(\chi_\perp\) separately, and classify the critical point without assuming they coincide. Route: Extension. Estimated time: 75 minutes. Prerequisite: Gaussian quadrature or Monte Carlo. Deliverable: two gain estimates, a pair trace, and a boundary statement. Hint: perturb the diagonal and off-diagonal kernel coordinates independently.
(Audit.) Paper audit: Locate a paper that uses “critical” and complete the Mathematical object, Resolution of the theory, Dynamic regime, Assumption stress test, Discriminating control, and Transfer verdict fields of the Paper Autopsy Protocol. Name its small parameter, retained order, and order of limits. As an agent-generated-proof control, remove one independence or width assumption silently, identify the first invalid step, and state the weakest repair. Route: Research. Estimated time: 90 minutes. Prerequisite: C11–C14. Deliverable: six-field autopsy plus repaired statement. Hint: “infinite width” is not an order of limits until depth is named.
Glorot, Xavier, and Yoshua Bengio. 2010. “Understanding the Difficulty of Training Deep Feedforward Neural Networks.”Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, 249–56. https://proceedings.mlr.press/v9/glorot10a.html.
He, Kaiming, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015. “Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification.”Proceedings of the IEEE International Conference on Computer Vision. https://doi.org/10.1109/ICCV.2015.123.
Roberts, Daniel A., and Sho Yaida. 2022. The Principles of Deep Learning Theory: An Effective Theory Approach to Understanding Neural Networks. Cambridge University Press. https://doi.org/10.1017/9781009023405.
Saxe, Andrew M., James L. McClelland, and Surya Ganguli. 2014. “Exact Solutions to the Nonlinear Dynamics of Learning in Deep Linear Neural Networks.”International Conference on Learning Representations. https://doi.org/10.48550/arXiv.1312.6120.
Schoenholz, Samuel S., Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein. 2017. “Deep Information Propagation.”International Conference on Learning Representations. https://openreview.net/forum?id=H1W1UN9gg.