Pith. sign in

REVIEW 2 cited by

Approximation and Estimation for High-Dimensional Deep Learning Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1809.03090 v2 pith:FB4S7V7U submitted 2018-09-10 stat.ML cs.LG

Approximation and Estimation for High-Dimensional Deep Learning Networks

classification stat.ML cs.LG
keywords networksparametersrisksamplesizeboundscontrolsdeep
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

It has been experimentally observed in recent years that multi-layer artificial neural networks have a surprising ability to generalize, even when trained with far more parameters than observations. Is there a theoretical basis for this? The best available bounds on their metric entropy and associated complexity measures are essentially linear in the number of parameters, which is inadequate to explain this phenomenon. Here we examine the statistical risk (mean squared predictive error) of multi-layer networks with $\ell^1$-type controls on their parameters and with ramp activation functions (also called lower-rectified linear units). In this setting, the risk is shown to be upper bounded by $[(L^3 \log d)/n]^{1/2}$, where $d$ is the input dimension to each layer, $L$ is the number of layers, and $n$ is the sample size. In this way, the input dimension can be much larger than the sample size and the estimator can still be accurate, provided the target function has such $\ell^1$ controls and that the sample size is at least moderately large compared to $L^3\log d$. The heart of the analysis is the development of a sampling strategy that demonstrates the accuracy of a sparse covering of deep ramp networks. Lower bounds show that the identified risk is close to being optimal.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Quantifying and Optimizing Simplicity via Polynomial Representations

    cs.AI 2026-05 unverdicted novelty 6.0

    Polynomial representations yield an effective-degree simplicity metric that predicts generalization across tasks and serves as a differentiable regularizer improving performance in classification and RL.

  2. Random Neural Network Expressivity for Non-Linear Partial Differential Equations

    math.NA 2026-05 unverdicted novelty 5.0

    Random neural networks achieve a dimension-free approximation rate of 1/2 for sufficiently regular time-dependent Sobolev functions and can efficiently approximate solutions to Porous Medium Equations and Compressible...