Pith. sign in

REVIEW 2 cited by

Representing Numbers in NLP: a Survey and a Vision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.13136 v1 pith:OYGGAZ2Z submitted 2021-03-24 cs.CL cs.AIcs.LG

Representing Numbers in NLP: a Survey and a Vision

classification cs.CL cs.AIcs.LG
keywords numbersnumeracyrepresentingtextvisionabstractalonganalyze
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

NLP systems rarely give special consideration to numbers found in text. This starkly contrasts with the consensus in neuroscience that, in the brain, numbers are represented differently from words. We arrange recent NLP work on numeracy into a comprehensive taxonomy of tasks and methods. We break down the subjective notion of numeracy into 7 subtasks, arranged along two dimensions: granularity (exact vs approximate) and units (abstract vs grounded). We analyze the myriad representational choices made by 18 previously published number encoders and decoders. We synthesize best practices for representing numbers in text and articulate a vision for holistic numeracy in NLP, comprised of design trade-offs and a unified evaluation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. FoNE: Precise Single-Token Number Embeddings via Fourier Features

    cs.CL 2025-02 unverdicted novelty 6.0

    FoNE encodes numbers as single tokens via Fourier features and outperforms subword and digit-wise embeddings on addition, subtraction, and multiplication with far less data.

  2. Predicting Post-Traumatic Epilepsy from Clinical Records using Large Language Model Embeddings

    cs.LG 2026-04 unverdicted novelty 5.0

    LLM embeddings from clinical records, fused with tabular data via gradient-boosted trees, predict post-traumatic epilepsy at AUC-ROC 0.892 and AUPRC 0.798.