Pith. sign in

REVIEW 2 cited by

Estimating the intrinsic dimension of datasets by a minimal neighborhood information

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1803.06992 v1 pith:P4U43EWZ submitted 2018-03-19 stat.ML cs.LG

Estimating the intrinsic dimension of datasets by a minimal neighborhood information

classification stat.ML cs.LG
keywords datamanifoldallowsanalysisblockdatasetsdimensiondistributed
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Analyzing large volumes of high-dimensional data is an issue of fundamental importance in data science, molecular simulations and beyond. Several approaches work on the assumption that the important content of a dataset belongs to a manifold whose Intrinsic Dimension (ID) is much lower than the crude large number of coordinates. Such manifold is generally twisted and curved, in addition points on it will be non-uniformly distributed: two factors that make the identification of the ID and its exploitation really hard. Here we propose a new ID estimator using only the distance of the first and the second nearest neighbor of each point in the sample. This extreme minimality enables us to reduce the effects of curvature, of density variation, and the resulting computational cost. The ID estimator is theoretically exact in uniformly distributed datasets, and provides consistent measures in general. When used in combination with block analysis, it allows discriminating the relevant dimensions as a function of the block size. This allows estimating the ID even when the data lie on a manifold perturbed by a high-dimensional noise, a situation often encountered in real world data sets. We demonstrate the usefulness of the approach on molecular simulations and image analysis.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ANN Search: Recall What Matters

    cs.IR 2026-06 conditional novelty 6.0

    ANN search quality is better assessed by 1/Ratio@k than Recall@k because the former tracks downstream task utility more closely while allowing substantially lower computational cost.

  2. MLLM-Microscope: Unlocking Hidden Structure Within Multimodal Large Language Models

    cs.CL 2026-05 unverdicted novelty 4.0

    MLLM-Microscope measures linearity, dimension and anisotropy of multimodal token streams in LLaVA-NeXT and OmniFusion, reporting high linearity overall and model-specific differences tied to modality fusion.