Pith. sign in
Pith Number

pith:BF2N5BFY

pith:2022:BF2N5BFYVYKEHSH2XCSV75QTDD
not attested not anchored not stored refs pending

Training Compute-Optimal Large Language Models

Aidan Clark, Arthur Mensch, Aurelia Guy, Bogdan Damoc, Diego de las Casas, Elena Buchatskaya, Eliza Rutherford, Erich Elsen, Eric Noland, George van den Driessche, Jack W. Rae, Johannes Welbl, Jordan Hoffmann, Karen Simonyan, Katie Millican, Laurent Sifre, Lisa Anne Hendricks, Oriol Vinyals, Sebastian Borgeaud, Simon Osindero, Tom Hennigan, Trevor Cai

For compute-optimal LLM training, scale model size and training tokens equally.

arxiv:2203.15556 v1 · 2022-03-29 · cs.CL · cs.LG

Add to your LaTeX paper
\usepackage{pith}
\pithnumber{BF2N5BFYVYKEHSH2XCSV75QTDD}

Prints a linked badge after your title and injects PDF metadata. Compiles on arXiv. Learn more · Embed verified badge

Record completeness

1 Bitcoin timestamp
2 Internet Archive
3 Author claim open · sign in to claim
4 Citations open
5 Replications open
Portable graph bundle live · download bundle · merged state
The bundle contains the canonical record plus signed events. A mirror can host it anywhere and recompute the same current state with the deterministic merge algorithm.

Claims

C1strongest claim

By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled.

C2weakest assumption

That the parametric scaling law fitted to models up to 16B parameters and 500B tokens extrapolates accurately to the 70B regime and that the assumed functional form of loss versus N and D is the correct one for deriving the optimum.

C3one line summary

For compute-optimal LLM training, model size and training tokens must be scaled at the same rate, as validated by Chinchilla (70B parameters, 4x data) outperforming larger models like Gopher (280B).

Formal links

2 machine-checked theorem links

Cited by

463 papers in Pith

Receipt and verification
First computed 2026-07-05T05:30:58.409780Z
Builder pith-number-builder-2026-05-17-v1
Signature Pith Ed25519 (pith-v1-2026-05) · public key
Schema pith-number/v1.0

Canonical hash

0974de84b8ae1443c8fab8a55ff61318e757277708d92ec739f3b51665861b86

Aliases

arxiv: 2203.15556 · arxiv_version: 2203.15556v1 · doi: 10.48550/arxiv.2203.15556 · pith_short_12: BF2N5BFYVYKE · pith_short_16: BF2N5BFYVYKEHSH2 · pith_short_8: BF2N5BFY
Agent API
Verify this Pith Number yourself
curl -sH 'Accept: application/ld+json' https://pith.science/pith/BF2N5BFYVYKEHSH2XCSV75QTDD \
  | jq -c '.canonical_record' \
  | python3 -c "import sys,json,hashlib; b=json.dumps(json.loads(sys.stdin.read()), sort_keys=True, separators=(',',':'), ensure_ascii=False).encode(); print(hashlib.sha256(b).hexdigest())"
# expect: 0974de84b8ae1443c8fab8a55ff61318e757277708d92ec739f3b51665861b86
Canonical record JSON
{
  "metadata": {
    "abstract_canon_sha256": "973f7852c6a4c37a6d3807cbd2333c34f6f3ab640a62a19cf0a866e0c3894233",
    "cross_cats_sorted": [
      "cs.LG"
    ],
    "license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
    "primary_cat": "cs.CL",
    "submitted_at": "2022-03-29T13:38:03Z",
    "title_canon_sha256": "cf1717178abca8c7941a23d05e156de6582b26e41269f4401fb09ae0d09b21ce"
  },
  "schema_version": "1.0",
  "source": {
    "id": "2203.15556",
    "kind": "arxiv",
    "version": 1
  }
}