REVIEW 2 major objections 2 minor 14 references
MLP residual networks selectively reduce the effective rank of the residual stream with depth only for short-correlation Markov inputs.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-27 13:40 UTC pith:GG2LXSX3
load-bearing objection The paper gives the first controlled quantitative test of the RG analogy in residual MLPs by tracking effective rank on Markov chains with known spectra, but does not check whether collapsed directions match the irrelevant input modes. the 2 major comments →
Rank Collapse, Fixed Points, and the Renormalization Group Structure of MLP Residual Networks
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
After training, the effective rank of the residual stream decreases monotonically with depth for chains whose correlation length is approximately 1 but remains unchanged for chains whose correlation length is approximately 7; the drop occurs position-wise and preserves precisely the degrees of freedom needed for masked prediction, while inter-layer kernel drift concentrates at one or two specific layer transitions.
What carries the argument
Effective rank of the residual stream, serving as an order parameter that registers progressive, selective integration of input features.
Load-bearing premise
The effective rank of the residual stream functions as a measurable RG order parameter whose behavior under controlled variation of input correlation length directly tests the coarse-graining interpretation.
What would settle it
If position-resolved effective rank shows no difference between short-correlation and long-correlation input chains after identical training, the selective-coarse-graining claim is falsified.
If this is right
- Networks on short-correlation sequences will exhibit depth-dependent rank reduction localized to task-irrelevant directions.
- Networks on long-correlation sequences will maintain full effective rank across all depths.
- Kernel matrices will exhibit drift only at one or two discrete layer transitions rather than uniformly.
- The directions that survive rank collapse will be exactly those required to solve the masked-prediction objective.
Where Pith is reading between the lines
- Similar rank trajectories may appear in language models when local predictability varies across a corpus.
- Intervening on the input spectrum before training should shift the depth at which rank collapse begins.
- The fixed-point plateaus imply that most layers after the first few act as identity maps once the relevant features are extracted.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies MLP residual networks trained on masked token prediction over synthetic Markov chain sequences with known spectral properties. It reports three empirical findings: (i) monotonic decrease in effective rank of the residual stream with depth after training; (ii) this rank collapse is selective, occurring for short-correlation-length chains (~1) but absent for long ones (~7), measured position-wise; (iii) inter-layer kernel drift concentrates at one or two transitions with the rest near fixed points. These are interpreted as the first quantitative, position-level evidence that such networks implement selective RG-style coarse-graining governed by the input distribution's spectral structure.
Significance. If the central claims hold, the work supplies the first controlled, quantitative test of the RG analogy in DNNs by defining an effective-rank order parameter, varying input correlation length, and verifying selectivity and fixed-point structure. This moves the analogy from qualitative to falsifiable and could inform both theoretical understanding of depth and practical architecture design.
major comments (2)
- [§4] §4 (Results on rank collapse): The selectivity claim—that the network discards precisely the irrelevant degrees of freedom per the RG relevance criterion—requires showing that the principal directions removed at each layer align with the eigenvectors of the input transition matrix having eigenvalues closest to zero. The reported monotonic rank reduction and correlation-length dependence establish depth-dependent compression but do not include this spectral decomposition or alignment test; without it, the observed collapse could reflect generic task-driven compression rather than spectral RG selectivity.
- [§3.2] §3.2 (Effective rank definition and measurement): The effective rank is positioned as the measurable RG order parameter, yet the manuscript provides no ablation or sensitivity analysis showing that alternative rank proxies (e.g., participation ratio vs. numerical rank at different thresholds) yield the same selective behavior under controlled input spectra. This choice is load-bearing for interpreting the order parameter as directly testing coarse-graining.
minor comments (2)
- The abstract states three findings but the main text should explicitly cross-reference the corresponding figures/tables for each (e.g., position-level rank curves for short vs. long correlation length).
- [§2] Notation for the Markov chain transition matrix and its eigenvalues should be introduced once in §2 and used consistently; current usage mixes spectral radius and correlation length without a single defining equation.
Simulated Author's Rebuttal
We thank the referee for the detailed and constructive report. We address each major comment below.
read point-by-point responses
-
Referee: [§4] §4 (Results on rank collapse): The selectivity claim—that the network discards precisely the irrelevant degrees of freedom per the RG relevance criterion—requires showing that the principal directions removed at each layer align with the eigenvectors of the input transition matrix having eigenvalues closest to zero. The reported monotonic rank reduction and correlation-length dependence establish depth-dependent compression but do not include this spectral decomposition or alignment test; without it, the observed collapse could reflect generic task-driven compression rather than spectral RG selectivity.
Authors: We agree that explicit eigenvector alignment would constitute stronger evidence. However, the experimental design varies the Markov chain correlation length, which directly sets the eigenvalue spectrum of the transition matrix (short length ~1 produces rapid eigenvalue decay away from 1; long length ~7 produces slower decay). The position-wise rank collapse occurs selectively only for the short-correlation case, indicating that the network compresses precisely the fast-decaying modes while preserving the slow ones required for the prediction task. This controlled spectral variation already tests the RG relevance criterion without needing per-layer eigenvector matching. We therefore do not plan to add the alignment analysis. revision: no
-
Referee: [§3.2] §3.2 (Effective rank definition and measurement): The effective rank is positioned as the measurable RG order parameter, yet the manuscript provides no ablation or sensitivity analysis showing that alternative rank proxies (e.g., participation ratio vs. numerical rank at different thresholds) yield the same selective behavior under controlled input spectra. This choice is load-bearing for interpreting the order parameter as directly testing coarse-graining.
Authors: We acknowledge the absence of such an ablation. The effective rank employed is the participation ratio, chosen for its continuity and direct sensitivity to the full spectrum. In the revised manuscript we will add a sensitivity check comparing this measure to numerical rank at thresholds 10^{-3} and 10^{-5}, confirming that the selective dependence on input correlation length is preserved across proxies. revision: yes
Circularity Check
Empirical measurements on synthetic data; no derivation reduces to fitted inputs or self-referential definitions
full rationale
The paper reports three empirical findings from training MLP residual stacks on masked prediction over Markov chains with controlled spectral properties: monotonic depth-dependent rank reduction, its selectivity for short vs. long correlation lengths at the position level, and concentration of kernel drift at specific transitions. These are direct observations of measured quantities (effective rank, inter-layer kernels) under explicit input variation; no equation or claim equates a 'prediction' to a fitted parameter by construction, invokes a self-citation as the sole justification for a uniqueness theorem, or renames a known pattern as a new RG derivation. The work is self-contained against external benchmarks (synthetic ground-truth spectra) and contains no load-bearing self-citation chain.
Axiom & Free-Parameter Ledger
axioms (1)
- domain assumption Effective rank of the residual stream constitutes a measurable RG order parameter
invented entities (1)
-
RG order parameter realized as effective rank
no independent evidence
read the original abstract
The analogy between deep neural network forward passes and renormalization group (RG) flows has been repeatedly noted in the literature, but existing treatments remain qualitative: depth is described as a coarse-graining scale, attention is likened to a partition function, and representations are said to flow toward fixed points. No existing work has defined a measurable RG order parameter, tested it under controlled variation of the input distribution, or made quantitative predictions that are empirically verified. We study the simplest architecture for which the analogy is tractable: a pure MLP residual stack trained on masked token prediction over synthetic Markov chain sequences with known spectral properties. We report three findings. (i) The effective rank of the residual stream decreases monotonically with depth after training, consistent with progressive integration of irrelevant degrees of freedom. (ii) This rank collapse is selective: it occurs for chains with short correlation length approximately 1 but is absent for chains with long correlation length approximately 7, measured at the position level to control for mean-pooling artifacts. The network preserves exactly the degrees of freedom relevant to the prediction task, the content of the RG relevance criterion. (iii) Inter-layer kernel drift is concentrated at one or two specific transitions, with the remainder of the network near a fixed point, consistent with a discrete fixed-point plateau. Together these findings constitute the first quantitative, position-level evidence that MLP residual networks implement a selective coarse-graining procedure governed by the spectral structure of the input distribution.
Figures
Reference graph
Works this paper leans on
-
[1]
Renormalization group and critical phenomena
Wilson, Kenneth G , journal=. Renormalization group and critical phenomena. 1971 , publisher=
1971
-
[2]
Physics Reports , volume=
The renormalization group and the epsilon expansion , author=. Physics Reports , volume=. 1974 , publisher=
1974
-
[3]
An exact mapping between the Variational Renormalization Group and Deep Learning
An exact mapping between the variational renormalization group and deep learning , author=. arXiv preprint arXiv:1410.3831 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[4]
Deep learning and the renormalization group
Renormalization group as a source of loss functions , author=. arXiv preprint arXiv:1301.3124 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[5]
Renormalization group for deep neural networks:
Bordelon, Blake and Pehlevan, Cengiz , journal=. Renormalization group for deep neural networks:
-
[6]
2026 , eprint=
Latent Object Permanence: Topological Phase Transitions, Free-Energy Principles, and Renormalization Group Flows in Deep Transformer Manifolds , author=. 2026 , eprint=
2026
-
[7]
Proceedings of the 36th International Conference on Machine Learning , pages=
Similarity of neural network representations revisited , author=. Proceedings of the 36th International Conference on Machine Learning , pages=
-
[8]
Tenney, Ian and Das, Dipanjan and Pavlick, Ellie , journal=
-
[9]
The effective rank:
Roy, Olivier and Vetterli, Martin , journal=. The effective rank:
-
[10]
Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina , journal=
-
[11]
2017 , eprint=
Adam: A Method for Stochastic Optimization , author=. 2017 , eprint=
2017
-
[12]
Gu, Albert and Dao, Tri , journal=. Mamba:
-
[13]
2026 , eprint=
A Theoretical Analysis of Mamba's Training Dynamics: Filtering Relevant Features for Generalization in State Space Models , author=. 2026 , eprint=
2026
-
[14]
2026 , eprint=
RGMem: Renormalization Group-inspired Memory Evolution for Language Agents , author=. 2026 , eprint=
2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.