Pith. sign in

REVIEW 5 cited by

Deep Generative Models through the Lens of the Manifold Hypothesis: A Survey and New Connections

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.02954 v2 pith:ZJMSWVCU submitted 2024-04-03 cs.LG cs.AIstat.ML

Deep Generative Models through the Lens of the Manifold Hypothesis: A Survey and New Connections

classification cs.LG cs.AIstat.ML
keywords dgmsmodelslensmanifoldgenerativeautoencodersdatadeep
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

In recent years there has been increased interest in understanding the interplay between deep generative models (DGMs) and the manifold hypothesis. Research in this area focuses on understanding the reasons why commonly-used DGMs succeed or fail at learning distributions supported on unknown low-dimensional manifolds, as well as developing new models explicitly designed to account for manifold-supported data. This manifold lens provides both clarity as to why some DGMs (e.g. diffusion models and some generative adversarial networks) empirically surpass others (e.g. likelihood-based models such as variational autoencoders, normalizing flows, or energy-based models) at sample generation, and guidance for devising more performant DGMs. We carry out the first survey of DGMs viewed through this lens, making two novel contributions along the way. First, we formally establish that numerical instability of likelihoods in high ambient dimensions is unavoidable when modelling data with low intrinsic dimension. We then show that DGMs on learned representations of autoencoders can be interpreted as approximately minimizing Wasserstein distance: this result, which applies to latent diffusion models, helps justify their outstanding empirical results. The manifold lens provides a rich perspective from which to understand DGMs, and we aim to make this perspective more accessible and widespread.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Localizing Memorized Regions in Diffusion Models via Coordinate-Wise Curvature Differences

    cs.LG 2026-05 unverdicted novelty 6.0

    Curvature-difference localization using underfitted baselines identifies overfitting-driven memorization in diffusion models and outperforms attention-based methods on Stable Diffusion.

  2. Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine

    cs.LG 2026-05 unverdicted novelty 6.0

    SiLD is a score-matching framework that learns both manifold projection and intrinsic density from a single objective, with proven sample complexity depending only on intrinsic dimension.

  3. What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion

    cs.CV 2026-05 unverdicted novelty 6.0

    Prior-Aligned AutoEncoders shape latent manifolds with spatial coherence, local continuity, and global semantics to improve latent diffusion, achieving SOTA gFID 1.03 on ImageNet 256x256 with up to 13x faster convergence.

  4. Bi-Lipschitz Autoencoder With Injectivity Guarantee

    cs.LG 2026-04 conditional novelty 6.0

    BLAE adds injective regularization via a separation criterion and bi-Lipschitz constraints to guarantee injectivity and geometric preservation in autoencoders, outperforming prior methods on manifold fidelity under sp...

  5. Why Code, Why Now: An Information-Theoretic Perspective on the Limits of Machine Learning

    cs.LG 2026-02 unverdicted novelty 6.0

    Task information structure determines ML scaling success, with code's dense verifiable signals enabling predictable progress while sparse-feedback tasks like typical RL do not.