Pith. sign in

REVIEW 2 major objections 2 minor 14 references

MLP residual networks selectively reduce the effective rank of the residual stream with depth only for short-correlation Markov inputs.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-27 13:40 UTC pith:GG2LXSX3

load-bearing objection The paper gives the first controlled quantitative test of the RG analogy in residual MLPs by tracking effective rank on Markov chains with known spectra, but does not check whether collapsed directions match the irrelevant input modes. the 2 major comments →

arxiv 2606.10324 v1 pith:GG2LXSX3 submitted 2026-06-09 cs.LG cond-mat.stat-mechstat.ML

Rank Collapse, Fixed Points, and the Renormalization Group Structure of MLP Residual Networks

classification cs.LG cond-mat.stat-mechstat.ML
keywords residual networksrank collapserenormalization groupMLPMarkov chainscoarse-grainingfixed pointseffective rank
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tests whether residual MLP stacks trained on masked token prediction implement a renormalization-group-style coarse-graining process. Using synthetic Markov chains whose spectral properties are known exactly, it tracks the effective rank of activations at every sequence position. Rank falls monotonically with depth when the input correlation length is short but stays flat when the correlation length is long, and the preserved directions match those required by the prediction task. Kernel matrices between layers stabilize after one or two transitions, leaving the rest of the network near a fixed point.

Core claim

After training, the effective rank of the residual stream decreases monotonically with depth for chains whose correlation length is approximately 1 but remains unchanged for chains whose correlation length is approximately 7; the drop occurs position-wise and preserves precisely the degrees of freedom needed for masked prediction, while inter-layer kernel drift concentrates at one or two specific layer transitions.

What carries the argument

Effective rank of the residual stream, serving as an order parameter that registers progressive, selective integration of input features.

Load-bearing premise

The effective rank of the residual stream functions as a measurable RG order parameter whose behavior under controlled variation of input correlation length directly tests the coarse-graining interpretation.

What would settle it

If position-resolved effective rank shows no difference between short-correlation and long-correlation input chains after identical training, the selective-coarse-graining claim is falsified.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Networks on short-correlation sequences will exhibit depth-dependent rank reduction localized to task-irrelevant directions.
  • Networks on long-correlation sequences will maintain full effective rank across all depths.
  • Kernel matrices will exhibit drift only at one or two discrete layer transitions rather than uniformly.
  • The directions that survive rank collapse will be exactly those required to solve the masked-prediction objective.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Similar rank trajectories may appear in language models when local predictability varies across a corpus.
  • Intervening on the input spectrum before training should shift the depth at which rank collapse begins.
  • The fixed-point plateaus imply that most layers after the first few act as identity maps once the relevant features are extracted.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper studies MLP residual networks trained on masked token prediction over synthetic Markov chain sequences with known spectral properties. It reports three empirical findings: (i) monotonic decrease in effective rank of the residual stream with depth after training; (ii) this rank collapse is selective, occurring for short-correlation-length chains (~1) but absent for long ones (~7), measured position-wise; (iii) inter-layer kernel drift concentrates at one or two transitions with the rest near fixed points. These are interpreted as the first quantitative, position-level evidence that such networks implement selective RG-style coarse-graining governed by the input distribution's spectral structure.

Significance. If the central claims hold, the work supplies the first controlled, quantitative test of the RG analogy in DNNs by defining an effective-rank order parameter, varying input correlation length, and verifying selectivity and fixed-point structure. This moves the analogy from qualitative to falsifiable and could inform both theoretical understanding of depth and practical architecture design.

major comments (2)
  1. [§4] §4 (Results on rank collapse): The selectivity claim—that the network discards precisely the irrelevant degrees of freedom per the RG relevance criterion—requires showing that the principal directions removed at each layer align with the eigenvectors of the input transition matrix having eigenvalues closest to zero. The reported monotonic rank reduction and correlation-length dependence establish depth-dependent compression but do not include this spectral decomposition or alignment test; without it, the observed collapse could reflect generic task-driven compression rather than spectral RG selectivity.
  2. [§3.2] §3.2 (Effective rank definition and measurement): The effective rank is positioned as the measurable RG order parameter, yet the manuscript provides no ablation or sensitivity analysis showing that alternative rank proxies (e.g., participation ratio vs. numerical rank at different thresholds) yield the same selective behavior under controlled input spectra. This choice is load-bearing for interpreting the order parameter as directly testing coarse-graining.
minor comments (2)
  1. The abstract states three findings but the main text should explicitly cross-reference the corresponding figures/tables for each (e.g., position-level rank curves for short vs. long correlation length).
  2. [§2] Notation for the Markov chain transition matrix and its eigenvalues should be introduced once in §2 and used consistently; current usage mixes spectral radius and correlation length without a single defining equation.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the detailed and constructive report. We address each major comment below.

read point-by-point responses
  1. Referee: [§4] §4 (Results on rank collapse): The selectivity claim—that the network discards precisely the irrelevant degrees of freedom per the RG relevance criterion—requires showing that the principal directions removed at each layer align with the eigenvectors of the input transition matrix having eigenvalues closest to zero. The reported monotonic rank reduction and correlation-length dependence establish depth-dependent compression but do not include this spectral decomposition or alignment test; without it, the observed collapse could reflect generic task-driven compression rather than spectral RG selectivity.

    Authors: We agree that explicit eigenvector alignment would constitute stronger evidence. However, the experimental design varies the Markov chain correlation length, which directly sets the eigenvalue spectrum of the transition matrix (short length ~1 produces rapid eigenvalue decay away from 1; long length ~7 produces slower decay). The position-wise rank collapse occurs selectively only for the short-correlation case, indicating that the network compresses precisely the fast-decaying modes while preserving the slow ones required for the prediction task. This controlled spectral variation already tests the RG relevance criterion without needing per-layer eigenvector matching. We therefore do not plan to add the alignment analysis. revision: no

  2. Referee: [§3.2] §3.2 (Effective rank definition and measurement): The effective rank is positioned as the measurable RG order parameter, yet the manuscript provides no ablation or sensitivity analysis showing that alternative rank proxies (e.g., participation ratio vs. numerical rank at different thresholds) yield the same selective behavior under controlled input spectra. This choice is load-bearing for interpreting the order parameter as directly testing coarse-graining.

    Authors: We acknowledge the absence of such an ablation. The effective rank employed is the participation ratio, chosen for its continuity and direct sensitivity to the full spectrum. In the revised manuscript we will add a sensitivity check comparing this measure to numerical rank at thresholds 10^{-3} and 10^{-5}, confirming that the selective dependence on input correlation length is preserved across proxies. revision: yes

Circularity Check

0 steps flagged

Empirical measurements on synthetic data; no derivation reduces to fitted inputs or self-referential definitions

full rationale

The paper reports three empirical findings from training MLP residual stacks on masked prediction over Markov chains with controlled spectral properties: monotonic depth-dependent rank reduction, its selectivity for short vs. long correlation lengths at the position level, and concentration of kernel drift at specific transitions. These are direct observations of measured quantities (effective rank, inter-layer kernels) under explicit input variation; no equation or claim equates a 'prediction' to a fitted parameter by construction, invokes a self-citation as the sole justification for a uniqueness theorem, or renames a known pattern as a new RG derivation. The work is self-contained against external benchmarks (synthetic ground-truth spectra) and contains no load-bearing self-citation chain.

Axiom & Free-Parameter Ledger

0 free parameters · 1 axioms · 1 invented entities

Only abstract available; ledger entries are inferred from the stated claims and limited to the minimal assumptions needed to interpret rank as an RG order parameter.

axioms (1)
  • domain assumption Effective rank of the residual stream constitutes a measurable RG order parameter
    Paper defines and uses rank collapse as the quantitative proxy for progressive integration of irrelevant degrees of freedom.
invented entities (1)
  • RG order parameter realized as effective rank no independent evidence
    purpose: To provide a measurable quantity that tracks selective coarse-graining in the network
    Introduced in the abstract as the central measurable quantity; no independent evidence outside the reported experiments is described.

pith-pipeline@v0.9.1-grok · 5801 in / 1328 out tokens · 20516 ms · 2026-06-27T13:40:24.450888+00:00 · methodology

0 comments
read the original abstract

The analogy between deep neural network forward passes and renormalization group (RG) flows has been repeatedly noted in the literature, but existing treatments remain qualitative: depth is described as a coarse-graining scale, attention is likened to a partition function, and representations are said to flow toward fixed points. No existing work has defined a measurable RG order parameter, tested it under controlled variation of the input distribution, or made quantitative predictions that are empirically verified. We study the simplest architecture for which the analogy is tractable: a pure MLP residual stack trained on masked token prediction over synthetic Markov chain sequences with known spectral properties. We report three findings. (i) The effective rank of the residual stream decreases monotonically with depth after training, consistent with progressive integration of irrelevant degrees of freedom. (ii) This rank collapse is selective: it occurs for chains with short correlation length approximately 1 but is absent for chains with long correlation length approximately 7, measured at the position level to control for mean-pooling artifacts. The network preserves exactly the degrees of freedom relevant to the prediction task, the content of the RG relevance criterion. (iii) Inter-layer kernel drift is concentrated at one or two specific transitions, with the remainder of the network near a fixed point, consistent with a discrete fixed-point plateau. Together these findings constitute the first quantitative, position-level evidence that MLP residual networks implement a selective coarse-graining procedure governed by the spectral structure of the input distribution.

Figures

Figures reproduced from arXiv: 2606.10324 by Irina Rish, Parviz Haggi-Mani.

Figure 1
Figure 1. Figure 1: Effective rank profiles across depth for the short-ξ chain (ξ = 1.2, flatten mode), at 11 training checkpoints from initialization (dark) to step 10,000 (light). Rank is near￾uniform at initialization and collapses mono￾tonically with depth after training, with a sharp fan visible between steps 8,000–9,000 at layers L3–L6. 0 2k 4k 6k 8k 10k Training step 0 2 4 6 8 10 12 14 16 Effective rank collapse window… view at source ↗
Figure 3
Figure 3. Figure 3: Effective rank at the final checkpoint under pool (dashed) and flatten (solid) [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Effective rank profiles across depth for the long- [PITH_FULL_IMAGE:figures/full_fig_p009_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Training loss curves for the short-ξ (blue, ξ = 1.20) and long-ξ (orange, ξ = 6.73) chains over 10,000 steps. Faint texture shows raw per-step loss; solid lines show exponential moving average (EMA-smoothed) loss. Dashed horizontal lines (left) mark H(π) = − P i πi log πi for each chain (H(π) = 2.722 for short-ξ; H(π) = 2.212 for long-ξ) — the lowest loss achievable by a position-wise model with no access … view at source ↗
Figure 6
Figure 6. Figure 6: Kernel drift profiles at the final checkpoint under pool (dashed) and flatten (solid) modes, short-ξ chain. Both modes show a roughly decreasing profile; pool is smoother. 0 1 1 2 2 3 3 4 4 5 5 6 Layer gap 0.0 0.1 0.2 0.3 0.4 0.5 Kernel drift (1 CKA) 0.56 0.38 Drift: long- ( = 6.73) mean-pool (n=512) flatten (n=16×seq) [PITH_FULL_IMAGE:figures/full_fig_p011_6.png] view at source ↗
Figure 8
Figure 8. Figure 8: Kernel drift profiles at the final checkpoint under pool and flatten modes. [PITH_FULL_IMAGE:figures/full_fig_p011_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Kernel drift heatmaps (flatten mode) across training. Rows = layer gaps; columns [PITH_FULL_IMAGE:figures/full_fig_p014_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

14 extracted references · 2 canonical work pages · 2 internal anchors

  1. [1]

    Renormalization group and critical phenomena

    Wilson, Kenneth G , journal=. Renormalization group and critical phenomena. 1971 , publisher=

  2. [2]

    Physics Reports , volume=

    The renormalization group and the epsilon expansion , author=. Physics Reports , volume=. 1974 , publisher=

  3. [3]

    An exact mapping between the Variational Renormalization Group and Deep Learning

    An exact mapping between the variational renormalization group and deep learning , author=. arXiv preprint arXiv:1410.3831 , year=

  4. [4]

    Deep learning and the renormalization group

    Renormalization group as a source of loss functions , author=. arXiv preprint arXiv:1301.3124 , year=

  5. [5]

    Renormalization group for deep neural networks:

    Bordelon, Blake and Pehlevan, Cengiz , journal=. Renormalization group for deep neural networks:

  6. [6]

    2026 , eprint=

    Latent Object Permanence: Topological Phase Transitions, Free-Energy Principles, and Renormalization Group Flows in Deep Transformer Manifolds , author=. 2026 , eprint=

  7. [7]

    Proceedings of the 36th International Conference on Machine Learning , pages=

    Similarity of neural network representations revisited , author=. Proceedings of the 36th International Conference on Machine Learning , pages=

  8. [8]

    Tenney, Ian and Das, Dipanjan and Pavlick, Ellie , journal=

  9. [9]

    The effective rank:

    Roy, Olivier and Vetterli, Martin , journal=. The effective rank:

  10. [10]

    Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina , journal=

  11. [11]

    2017 , eprint=

    Adam: A Method for Stochastic Optimization , author=. 2017 , eprint=

  12. [12]

    Gu, Albert and Dao, Tri , journal=. Mamba:

  13. [13]

    2026 , eprint=

    A Theoretical Analysis of Mamba's Training Dynamics: Filtering Relevant Features for Generalization in State Space Models , author=. 2026 , eprint=

  14. [14]

    2026 , eprint=

    RGMem: Renormalization Group-inspired Memory Evolution for Language Agents , author=. 2026 , eprint=