Pith. sign in

Paper Citation Record · LEDGER

Training Transformers for KV Cache Compressibility

As of 23 July 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2605.05971.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.05971 v2

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T06:01:40.766843Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-20T06:30:07.809122+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact23
  • verified fuzzy35
  • unresolved3
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1cbf67cf-30da-48db-9871-2b97be1dc410 · outbound

This paper cites Can Foundation Models Help Us Achieve Perfect Secrecy?.

Training Transformers for KV Cache Compressibility Can Foundation Models Help Us Achieve Perfect Secrecy?

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.987477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:b70001ca5f29b24fd99308abbcd9898e36e3e54496493a56d17f1eacfe50cac3

Observation aef4b139-e3c5-4177-9b93-3923e41c6c8d · outbound

This paper cites Longbench v2: Towards deeper understanding and reason- ing on realistic long-context multitasks.

Training Transformers for KV Cache Compressibility Longbench v2: Towards deeper understanding and reason- ing on realistic long-context multitasks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.294421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:bea46f598b1d82e3e01337b922f83786ce6cda49d7a5e0cee519794eb161533c

Observation b4d5c7da-b73d-4ec3-a51a-c821f65cb593 · outbound

This paper cites Longformer: The Long-Document Transformer.

Training Transformers for KV Cache Compressibility Longformer: The Long-Document Transformer

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:21.990094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:2ab951d80ebd3e2e45a0fb26e5fddab25d3c265cf2a19d5475d2805300c4863a

Observation 3493bf0e-5511-4291-a28e-3e97eaa32b7a · outbound

This paper cites PIQA: Reasoning about physical commonsense in natural language.

Training Transformers for KV Cache Compressibility PIQA: Reasoning about physical commonsense in natural language

Reference 4

Resolution
verified exact
doi, observed 2026-05-13T06:02:21.793734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:402771b890e019764270f423e6c5f4cb970234e4a4988f238b93c4bed2403278

Observation 7be11d93-61c2-4889-b04b-0795f160ba40 · outbound

This paper cites an unresolved cited work.

Training Transformers for KV Cache Compressibility Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-13T06:02:24.289944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:2b0d20bd90e60eb2b575fb56076d50f105881641d9694500a7845dad309a3823

Observation 9936c8c6-164c-45d0-93f7-f81526078cb2 · outbound

This paper cites PyramidKV: Dynamic kv cache compression based on pyramidal information funneling.

Training Transformers for KV Cache Compressibility PyramidKV: Dynamic kv cache compression based on pyramidal information funneling

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.287922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:f0939a05e5e9c74903d1abed40d78e8b619e23ea0c074ddf1349cdc8a34fe083

Observation 66d1b34d-c5f2-42f7-8b5e-bc7920e8e19c · outbound

This paper cites Doc-to-lora: Learning to instantly internalize contexts.

Training Transformers for KV Cache Compressibility Doc-to-lora: Learning to instantly internalize contexts

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.959138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:e981a7b8650e50c0f0e354c1dd3125c2615f41d7d3f2dfb1cb463213c16ce914

Observation a658111a-bcc5-49ea-9152-d1f6504481de · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Training Transformers for KV Cache Compressibility Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:21.981495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:99946ea3855ba121d7d8cc7178d004c57f75104ca95e8ed1a8433d8400fb40fd

Observation 1eba9c7d-de47-40c9-8bcf-90e758f08cbd · outbound

This paper cites Generative Adapter: Contextualizing Language Models in Parameters with A Single Forward Pass.

Training Transformers for KV Cache Compressibility Generative Adapter: Contextualizing Language Models in Parameters with A Single Forward Pass

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.953195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:379e610cad83e3a75054980f2d8764bc2b2c3872012142c6564d2e2b3fa7a977

Observation 3c77b851-8b43-40e8-9fbd-2e4dda0c619b · outbound

This paper cites Adapting language models to compress contexts.

Training Transformers for KV Cache Compressibility Adapting language models to compress contexts

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.285790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:a04df0e40a66ee29bb2094163af7612c8988dc671336a769e0410ce295adfe5b

Observation d9c843de-99c1-4fab-9c12-28c8a18157f2 · outbound

This paper cites Rethinking Attention with Performers.

Training Transformers for KV Cache Compressibility Rethinking Attention with Performers

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:21.943990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:28d38425f0d123d59e4ed384f97f3718a1d51c9cb1178800d87a600b5e4fc48d

Observation 4db62a8c-2dd1-455a-a48d-c04e2aa53630 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Training Transformers for KV Cache Compressibility Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:21.967996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:670b633dd811071779da7c5de477ff2a4ce7fa523f6dc6176926039ebc885613

Observation f93f265d-ef37-4145-ac87-3a3901e4034c · outbound

This paper cites InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers).

Training Transformers for KV Cache Compressibility InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers)

Reference 13

Resolution
metadata mismatch
doi, observed 2026-05-13T06:02:21.797274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:7ac12fe588d4a22488a147e79e0cedd0c1ae643cef9a154c7167e167bd703174

Observation f84f24d0-bfe3-49f4-9a63-c1def112e4f3 · outbound

This paper cites Approximation by superpositions of a sigmoidal function.Mathematics of control, signals and systems, 2(4):303–314.

Training Transformers for KV Cache Compressibility Approximation by superpositions of a sigmoidal function.Mathematics of control, signals and systems, 2(4):303–314

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.278883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:ae16157ab2cffbdefe687cf7f72da2d8d6cb1eaa5653d91687ebefde24551838

Observation dae1ed25-d9e0-412f-9bfd-c30c20856816 · outbound

This paper cites The centered convex body whose marginals have the heaviest tails.

Training Transformers for KV Cache Compressibility The centered convex body whose marginals have the heaviest tails

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.946969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:2112604c700b0152484bb79cf5753c23c3046d22b3acbb4a8d5551ddfa67e575

Observation 9e953d73-6031-42fc-9a30-726e1da7f055 · outbound

This paper cites Cartridges: Lightweight and general-purpose long context representations via self-study.

Training Transformers for KV Cache Compressibility Cartridges: Lightweight and general-purpose long context representations via self-study

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.283639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:4dcdaa56d09c1ec904a01474a9dc6327fbb2ed2fb7ea160dfc04c1d3424a93d4

Observation b5a6d708-345a-4c78-bccf-c14a19e9d834 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Training Transformers for KV Cache Compressibility Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:21.997183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:0c74eba1c14f5404cfd8d8329989525edda933941ac56b2a0cbfc9c84a74b977

Observation fa307a39-d96e-414a-895a-048eb46a3527 · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Training Transformers for KV Cache Compressibility Efficiently Modeling Long Sequences with Structured State Spaces

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:21.984721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:4f13ce29c8f12d9754ad46f9b0c0e7fa73aaa33e73a17db8fe24c196b6e5eb04

Observation af3ae607-84bb-4fd7-bb85-7ca3fc615b42 · outbound

This paper cites Lighte- val: A lightweight framework for llm evaluation.

Training Transformers for KV Cache Compressibility Lighte- val: A lightweight framework for llm evaluation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.281046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:7fd9ef0808ac2b242584292a10789faf06fe08a45b373dabb83ba98af2cad571

Observation 90d9039f-cf5f-484a-8221-1c221834db3b · outbound

This paper cites Delta-net: Real-time network verification using atoms.

Training Transformers for KV Cache Compressibility Delta-net: Real-time network verification using atoms

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.296958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:96a037b69034064696dceb8aea95cbe45a7a0c77282e8362709475fdc3047bdc

Observation 7eec02fd-e121-411b-a8aa-89dbac8dd5cb · outbound

This paper cites Approximation capabilities of multilayer feedforward networks.Neural networks, 4(2):251–257.

Training Transformers for KV Cache Compressibility Approximation capabilities of multilayer feedforward networks.Neural networks, 4(2):251–257

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.274052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:7b5046a25980227e6d87ac5732ad4478b6a7ffd545fec28d43ef03507ccf8d76

Observation 68a00469-4c6f-486a-9f23-033449813713 · outbound

This paper cites Dynamic Chunking for End-to-End Hierarchical Sequence Modeling.

Training Transformers for KV Cache Compressibility Dynamic Chunking for End-to-End Hierarchical Sequence Modeling

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:02:21.965287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:62db94a122e88db83fca275ee1930001d1912ce013cf514e0ba58f557f6d1ef0

Observation 9a326a7c-54ae-4f48-bf4c-838d0e8a0510 · outbound

This paper cites FinanceBench: A New Benchmark for Financial Question Answering.

Training Transformers for KV Cache Compressibility FinanceBench: A New Benchmark for Financial Question Answering

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:01:13.018607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:86fdc5048a771da2fb32eb2cda44edbe727ce0e0f397bc95955be3323195a27c

Observation 452ddc86-650e-4627-9d2e-738e064716b8 · outbound

This paper cites LLMLingua: Com- pressing prompts for accelerated inference of large language models.

Training Transformers for KV Cache Compressibility LLMLingua: Com- pressing prompts for accelerated inference of large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.271575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:37ca7a4177a7bcaf9119a753c6325f2549d2e5f567ec1abb060cd7ea1419a379

Observation 147a14f9-4813-44e6-8fc8-36fa5f80ccbd · outbound

This paper cites LongLLMLingua: Accelerating and enhancing llms in long context scenarios via prompt compression.

Training Transformers for KV Cache Compressibility LongLLMLingua: Accelerating and enhancing llms in long context scenarios via prompt compression

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.276579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:55a318d312552cd31b9ac0de676a4bdbb0d69623f64c860dcdebd122d6127027

Observation 7b6c23f3-f5ff-4d25-a284-f34b0ee6cc09 · outbound

This paper cites Optimal experimental designs.The Annals of Mathemat- ical Statistics, 37(4):783–815.

Training Transformers for KV Cache Compressibility Optimal experimental designs.The Annals of Mathemat- ical Statistics, 37(4):783–815

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.269453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:edd850c79b02ce930967697e3e9544843e80feb97f13605b5bce2deaf7dcffcb

Observation 7677d385-0243-47cc-85d4-c0dc0da6fdd1 · outbound

This paper cites Tchebycheff systems: With applications in analysis and statistics.(No Title).

Training Transformers for KV Cache Compressibility Tchebycheff systems: With applications in analysis and statistics.(No Title)

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.264839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:72dadb1b00522bfd299bd386d5fa0c620d5a5df7bbc9626816653ffe4500adfa

Observation 14256ac9-4194-45c5-81a1-37e532fa0702 · outbound

This paper cites Chebyshevian spline functions.Siam Journal on Numerical Analysis, 3(3):514–543.

Training Transformers for KV Cache Compressibility Chebyshevian spline functions.Siam Journal on Numerical Analysis, 3(3):514–543

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.267004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:14ea3c769900b3b4654ee841ca4c8c373fefbb1ca3e2e576543dde6c226caa88

Observation dc366d46-17fc-4c89-9a35-3a056296534a · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

Training Transformers for KV Cache Compressibility Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.260987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:8ff44c2e814fc01f23d23ef92f479a6b30de1ba495614e12ff00bcbeeba69bbd

Observation 265f9aee-8bae-4477-8744-baafd32ad3aa · outbound

This paper cites Kvzip: Query-agnostic kv cache compression with context reconstruction.

Training Transformers for KV Cache Compressibility Kvzip: Query-agnostic kv cache compression with context reconstruction

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.971227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:0bf9a161250610bf91dc6d228b80173910bcb63314be7f386c6ab94bd7a9a760

Observation 0c3ab234-0bc3-47c1-a830-35e95aba6680 · outbound

This paper cites Lexico: Extreme KV cache compression via sparse coding over universal dictionaries.

Training Transformers for KV Cache Compressibility Lexico: Extreme KV cache compression via sparse coding over universal dictionaries

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.254451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:ecd4bf7fa08c55527862e5b2c6c6c13298bc2e8b9f5805c795be840bf552796a

Observation d9c0c9a0-9b89-4abd-a86b-dc07dc7fb28c · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive NLP tasks.

Training Transformers for KV Cache Compressibility Retrieval-augmented generation for knowledge-intensive NLP tasks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.256522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:a7913451e76aca66f76e355259f29a0820f56518fee93f25fb4b345f5bd9d4db

Observation b36181d0-9a4f-491d-9dbb-4544c41064c2 · outbound

This paper cites Compressing context to enhance inference efficiency of large language models.

Training Transformers for KV Cache Compressibility Compressing context to enhance inference efficiency of large language models

Reference 33

Resolution
malformed identifier
raw_fallback, observed 2026-05-13T06:02:24.258780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:c1852883406da307e344da0c84ed80ab3cb646525fd4f72b74dcbf828b8573ef

Observation 0d344f97-1b95-470f-8156-b066ee485948 · outbound

This paper cites SnapKV: Llm knows what you are looking for before generation.

Training Transformers for KV Cache Compressibility SnapKV: Llm knows what you are looking for before generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.262982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:f8c847a5c88924133fecf9b0cc4f40ac16fcc23b527dac38e9af27f43e63df3b

Observation ced269bf-c1b2-4443-959a-8fa0189b7ab3 · outbound

This paper cites LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence.

Training Transformers for KV Cache Compressibility LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.993912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:353063194465e11990bdd1c00b6a0dcb7447302434c5dad9fd248cb58ca316a8

Observation 89401e05-302f-4574-9d6f-221366601686 · outbound

This paper cites SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass.

Training Transformers for KV Cache Compressibility SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-21T03:03:41.308564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:92b436c0b876e4da83be3d07396dba8315101a2c31699f219cd975514df45d28

Observation 628d164d-0462-476f-9b91-c95e952b6feb · outbound

This paper cites Pointer sentinel mixture models.

Training Transformers for KV Cache Compressibility Pointer sentinel mixture models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.292079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:afa1ca3e9854d5765aee16f79374c46143a3d4aefa8108cfe1fdbf9f0745f85a

Observation 8976502a-1931-45cc-a261-2796fd872c02 · outbound

This paper cites an unresolved cited work.

Training Transformers for KV Cache Compressibility Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-13T06:02:24.250037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:10a6eda5198215eb5610e3f67182f189f23590e04b053f5e6090232d7218c18e

Observation bd3951d8-cdd7-41c6-9b01-43f233d1d814 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Training Transformers for KV Cache Compressibility Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 39

Resolution
verified exact
doi, observed 2026-05-13T06:02:21.790714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:ae68223993f156f55eb89fdfa2c1122d6c08f657fd2e4ae80a53e665246a0158

Observation e313e9fa-7d48-4e25-b6fe-6b3a50bef301 · outbound

This paper cites Learning to compress prompts with gist tokens.

Training Transformers for KV Cache Compressibility Learning to compress prompts with gist tokens

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.235100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:32d3548c3907353eeca390306b40a01090f8e5cd690f533696f3f7742bc7d785

Observation e5718b8d-1ccb-478e-8d9b-d9a69704d022 · outbound

This paper cites Using an llm to help with code understanding.

Training Transformers for KV Cache Compressibility Using an llm to help with code understanding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.220818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:6a1a8c978f59748ff5650ded2421a7cfd89898bf1f99876a36760c959b7ea4ff

Observation c076a8b2-29b6-4f6a-a674-b13ac4d9bd53 · outbound

This paper cites Context Engineering - Short-Term Memory Management with Sessions from OpenAI Agents SDK, September 2025.

Training Transformers for KV Cache Compressibility Context Engineering - Short-Term Memory Management with Sessions from OpenAI Agents SDK, September 2025

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.223236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:8a90b4e2e50d0a0597793cf7f3ed137b366d42e4b8500e48bfa296d1bb1a3021

Observation 436b9803-9b50-47e7-8d5a-f49bbe585cbc · outbound

This paper cites Transformers are multi-state RNNs.

Training Transformers for KV Cache Compressibility Transformers are multi-state RNNs

Reference 43

Resolution
verified exact
doi, observed 2026-05-13T06:02:21.803106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:2dec9b067c06f87b1fcf285f7f60eb5b66b4f2414e1281fa6dc891beeedc14ee

Observation 393bdcbc-4be4-4d49-99d4-95bf24436637 · outbound

This paper cites The fineweb datasets: Decanting the web for the finest text data at scale.Advances in Neural Information Processing Systems, 37:30811–30849.

Training Transformers for KV Cache Compressibility The fineweb datasets: Decanting the web for the finest text data at scale.Advances in Neural Information Processing Systems, 37:30811–30849

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.227937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:9b56f64d601e0a220483a7d0a2e9484e89e37db4085674f15bb0310c4ea229dc

Observation 4c76dc23-6216-4dfa-b833-1a6b6a2f153f · outbound

This paper cites Compressive Transformers for Long-Range Sequence Modelling.

Training Transformers for KV Cache Compressibility Compressive Transformers for Long-Range Sequence Modelling

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:46:17.004740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:73f07ca41a606dec35f97864c57e1bf1cafc9a553e720e73205989316e5f2702

Observation 40abb738-a348-46b4-aa94-045157a1adbb · outbound

This paper cites Effective context engi- neering for ai agents, September 2025.

Training Transformers for KV Cache Compressibility Effective context engi- neering for ai agents, September 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.218665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:5783dd6649eff81637c567b60ec35cf96014accb812b813b1509a9aa2732c8b3

Observation 59cfac00-a8e9-4f9b-a166-4993e0673cf3 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , author=.

Training Transformers for KV Cache Compressibility Proceedings of the AAAI Conference on Artificial Intelligence , author=

Reference 47

Resolution
verified exact
doi, observed 2026-05-13T06:02:21.786005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:c2f2f49bdb7bff288271993e45bd527ad7a12ed198be195617fc2ddbba8151a9

Observation 1d3aa75c-7184-4ffb-8779-6622470094bd · outbound

This paper cites Representational strengths and limitations of transformers.Advances in Neural Information Processing Systems, 36:36677–36707.

Training Transformers for KV Cache Compressibility Representational strengths and limitations of transformers.Advances in Neural Information Processing Systems, 36:36677–36707

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.213905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:3c197eaccf8bd543a3549981654a093abc0a9792e508d672cc1907d505cd7bc9

Observation fe83771e-4d4a-471c-89e7-526400e0d2da · outbound

This paper cites Social IQa: Commonsense reasoning about social interactions.

Training Transformers for KV Cache Compressibility Social IQa: Commonsense reasoning about social interactions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.225530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:a1927c1977b3de2ac10515d571d773e1e01606d3b7c2a54902fa05ee99e09c06

Observation 3f357129-8531-4ade-8248-1d2c178f607b · outbound

This paper cites Social IQ a: Commonsense reasoning about social interactions.

Training Transformers for KV Cache Compressibility Social IQ a: Commonsense reasoning about social interactions

Reference 50

Resolution
metadata mismatch
doi, observed 2026-05-13T06:02:21.800181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:bf8e05ab689f37026891df28e1d67d211f46453ad03ce74017cdeeeaf53ca01e

Observation ce7ad80b-0f42-46bf-8ae6-71a8d92048bd · outbound

This paper cites QUEST: Query-aware sparsity for efficient long-context LLM inference.

Training Transformers for KV Cache Compressibility QUEST: Query-aware sparsity for efficient long-context LLM inference

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.252230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:a0c4bf7b8261479fd49cf5f326f1a32dff66078e7610f1d149c58fcafdaf06e7

Observation 3080a2c0-cb92-4ffd-81ed-a3fcae3bda5f · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Training Transformers for KV Cache Compressibility Qwen2.5: A party of foundation models, September 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.204011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:4cbff4da9b7d5fdf1f5e911baa0b51d95031ad68f42ae2b6352b4d70e0826956

Observation 8a2a4e3b-5597-493d-bd83-ac4725c3f36a · outbound

This paper cites Efficient streaming language models with attention sinks.

Training Transformers for KV Cache Compressibility Efficient streaming language models with attention sinks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.208978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:01bf26230e0228629df988b63aec4b395d399c0bf51cb62e444799b4981e9b4d

Observation 0bd4eb5c-44b2-47f3-b70f-859d0874db08 · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

Training Transformers for KV Cache Compressibility DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:49:16.947404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:a836cb7a0366ce9891cc0cbfaa94bb7c2eddae7d94eebab71f9790fffc63ac35

Observation 6113bfcc-ac39-448c-a9d4-21b8a1d1bba4 · outbound

This paper cites Gated Delta Networks: Improving Mamba2 with Delta Rule.

Training Transformers for KV Cache Compressibility Gated Delta Networks: Improving Mamba2 with Delta Rule

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:50:25.025022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:61b2649f7e1b61150a3fcd6cd4bf28c7649c56ab2908d4318d9a5c3056884c92

Observation dc536cf1-3d8d-4043-b146-f4033d7e12e7 · outbound

This paper cites arXiv preprint arXiv:2407.15160 , year =.

Training Transformers for KV Cache Compressibility arXiv preprint arXiv:2407.15160 , year =

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:22.000080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:8e62c7fb778eb0e40f0cf3a7119a887a971d6aed534100499114a0504db53b8b

Observation 2bb92f07-362e-4e28-b0da-f1f7bf28eabe · outbound

This paper cites Deep sets.Advances in neural information processing systems, 30.

Training Transformers for KV Cache Compressibility Deep sets.Advances in neural information processing systems, 30

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.237288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:da2e17b99d65b00c74c7cc2c5390527df0c44d805f29f9ec7bb2a86736a4bfff

Observation 5cffa41d-2c43-4d58-931e-df3ed14a0102 · outbound

This paper cites Big bird: Transformers for longer sequences.Advances in neural information processing systems, 33: 17283–17297.

Training Transformers for KV Cache Compressibility Big bird: Transformers for longer sequences.Advances in neural information processing systems, 33: 17283–17297

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.230077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:ebc13b2f3a379658bf22ab813c77337619bb3f326db9d917d6f8c30f8673bc62

Observation 6557247c-93a1-4764-9f98-24dfbadeb32a · outbound

This paper cites URL https:// doi.org/10.18653/v1/p19-1472.

Training Transformers for KV Cache Compressibility URL https:// doi.org/10.18653/v1/p19-1472

Reference 59

Resolution
verified exact
doi, observed 2026-05-13T06:02:21.806038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:8c80b80c0957e3634106724289d5b24a5469c90577d182b46a7602b8acb0dcad

Observation 78c43fa5-9a10-49a3-84fc-d3d4cf794a20 · outbound

This paper cites H2O: Heavy-hitter oracle for efficient generative inference of large language models.

Training Transformers for KV Cache Compressibility H2O: Heavy-hitter oracle for efficient generative inference of large language models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.216232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:40e849529ef431413935c25e94c52566d3efe0135703fbcefa56983ea333fe6c

Observation 7444f49f-cb4e-4adb-83e3-5f870f1a0af6 · outbound

This paper cites Lifelong learning of large language model based agents: A roadmap.IEEE Transactions on Pattern Analysis and Machine Intelligence.

Training Transformers for KV Cache Compressibility Lifelong learning of large language model based agents: A roadmap.IEEE Transactions on Pattern Analysis and Machine Intelligence

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.232644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:6b2584971a073dbdc4ca6a81415341ccc7f9b04bb309dfc7b55685335c7d61d0

Observation 7b4e5337-878f-4ec0-a002-6601849bdb1c · outbound

This paper cites Fast kv compaction via attention matching.

Training Transformers for KV Cache Compressibility Fast kv compaction via attention matching

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.244984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:cb97b345f67e20c458693233bb7b515f993704fd863bddcdb8f1c00a6f506a78

Observation c6758ec7-0854-4dcb-8ee1-936b6165bf6a · outbound

This paper cites an)∈A n withn≤N, ∥f(a)−M(a)∥< ε.(70).

Training Transformers for KV Cache Compressibility an)∈A n withn≤N, ∥f(a)−M(a)∥< ε.(70)

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.247803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:e61b7d7b0ef46d4a88cbfaa9047b611e8fd1c44af3ef8f6a8dd050fcf98cbc5a

Observation 65356efd-af18-492c-b640-be94d32de4c7 · outbound

This paper cites an unresolved cited work.

Training Transformers for KV Cache Compressibility Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-13T06:02:24.206400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:166872d73939eca6b87c5c162021b549c95b518275376852413ea6ebff57091d

Observation 7ac604d7-f5e5-41b1-a5f0-4b3f2221255f · outbound

This paper cites an)∈A n withn≤N, ∥f(a)−M(a)∥< ε.(97).

Training Transformers for KV Cache Compressibility an)∈A n withn≤N, ∥f(a)−M(a)∥< ε.(97)

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.211502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:423a3599731012d01f63e4648e5a87fc6132903fb13d00b15ae02ac04983fe51

Observation a0c110ff-ee9b-45a9-8cad-83a1c8769821 · outbound

This paper cites By the same argument as in the proof of Lemma A.7, there exist functions ϕ:R d0 →R d1 andρ:R d1 →R dout such that for everya= (a 1.

Training Transformers for KV Cache Compressibility By the same argument as in the proof of Lemma A.7, there exist functions ϕ:R d0 →R d1 andρ:R d1 →R dout such that for everya= (a 1

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.950355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:284045c8c400b9b9366b31b9bbe3956ceded844c70fe8d7c36a8b0dd70b6a392

Pith citing papers

No inbound Pith citation observations are available.