{"paper":{"title":"Training Compute-Optimal Large Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"For compute-optimal LLM training, scale model size and training tokens equally.","cross_cats":["cs.LG"],"primary_cat":"cs.CL","authors_text":"Aidan Clark, Arthur Mensch, Aurelia Guy, Bogdan Damoc, Diego de las Casas, Elena Buchatskaya, Eliza Rutherford, Erich Elsen, Eric Noland, George van den Driessche, Jack W. Rae, Johannes Welbl, Jordan Hoffmann, Karen Simonyan, Katie Millican, Laurent Sifre, Lisa Anne Hendricks, Oriol Vinyals, Sebastian Borgeaud, Simon Osindero, Tom Hennigan, Trevor Cai","submitted_at":"2022-03-29T13:38:03Z","abstract_excerpt":"We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant. By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the parametric scaling law fitted to models up to 16B parameters and 500B tokens extrapolates accurately to the 70B regime and that the assumed functional form of loss versus N and D is the correct one for deriving the optimum.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"For compute-optimal LLM training, model size and training tokens must be scaled at the same rate, as validated by Chinchilla (70B parameters, 4x data) outperforming larger models like Gopher (280B).","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"For compute-optimal LLM training, scale model size and training tokens equally.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"06a57e72b7cd06f29a99428ec39c98c3c10a93004cb2bf31cc8002f4e72d367e"},"source":{"id":"2203.15556","kind":"arxiv","version":1},"verdict":{"id":"02bff6e5-f4fc-49ef-a2c4-ab3050de5241","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-10T15:55:30.872186Z","strongest_claim":"By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled.","one_line_summary":"For compute-optimal LLM training, model size and training tokens must be scaled at the same rate, as validated by Chinchilla (70B parameters, 4x data) outperforming larger models like Gopher (280B).","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the parametric scaling law fitted to models up to 16B parameters and 500B tokens extrapolates accurately to the 70B regime and that the assumed functional form of loss versus N and D is the correct one for deriving the optimum.","pith_extraction_headline":"For compute-optimal LLM training, scale model size and training tokens equally."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2203.15556/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":2,"snapshot_sha256":"f47260aa49859345287a9789a99e3f2ace0cf8fa7c0d168c7902ec10820bdf9b"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}