pith:BF2N5BFY
Training Compute-Optimal Large Language Models
For compute-optimal LLM training, scale model size and training tokens equally.
arxiv:2203.15556 v1 · 2022-03-29 · cs.CL · cs.LG
Add to your LaTeX paper
\usepackage{pith}
\pithnumber{BF2N5BFYVYKEHSH2XCSV75QTDD}
Prints a linked badge after your title and injects PDF metadata. Compiles on arXiv. Learn more · Embed verified badge
Record completeness
Claims
By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled.
That the parametric scaling law fitted to models up to 16B parameters and 500B tokens extrapolates accurately to the 70B regime and that the assumed functional form of loss versus N and D is the correct one for deriving the optimum.
For compute-optimal LLM training, model size and training tokens must be scaled at the same rate, as validated by Chinchilla (70B parameters, 4x data) outperforming larger models like Gopher (280B).
Formal links
Cited by
Receipt and verification
| First computed | 2026-07-05T05:30:58.409780Z |
|---|---|
| Builder | pith-number-builder-2026-05-17-v1 |
| Signature | Pith Ed25519
(pith-v1-2026-05) · public key |
| Schema | pith-number/v1.0 |
Canonical hash
0974de84b8ae1443c8fab8a55ff61318e757277708d92ec739f3b51665861b86
Aliases
· · · · ·Agent API
Verify this Pith Number yourself
curl -sH 'Accept: application/ld+json' https://pith.science/pith/BF2N5BFYVYKEHSH2XCSV75QTDD \
| jq -c '.canonical_record' \
| python3 -c "import sys,json,hashlib; b=json.dumps(json.loads(sys.stdin.read()), sort_keys=True, separators=(',',':'), ensure_ascii=False).encode(); print(hashlib.sha256(b).hexdigest())"
# expect: 0974de84b8ae1443c8fab8a55ff61318e757277708d92ec739f3b51665861b86
Canonical record JSON
{
"metadata": {
"abstract_canon_sha256": "973f7852c6a4c37a6d3807cbd2333c34f6f3ab640a62a19cf0a866e0c3894233",
"cross_cats_sorted": [
"cs.LG"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"primary_cat": "cs.CL",
"submitted_at": "2022-03-29T13:38:03Z",
"title_canon_sha256": "cf1717178abca8c7941a23d05e156de6582b26e41269f4401fb09ae0d09b21ce"
},
"schema_version": "1.0",
"source": {
"id": "2203.15556",
"kind": "arxiv",
"version": 1
}
}