Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T16:57:56.333595Z
Paper Citation Record · LEDGER
As of 23 July 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2606.01503.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T16:57:56.333595Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-20T06:30:07.809122+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bbfd7d18-a573-4d64-a687-e89dd65c2022 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Token merging: Your vit but faster, 2023
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5182b6d0-8fae-42f7-8347-08ec0a408930 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Sharegpt4v: Improving large multi-modal models with better captions
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faa7894c-0022-48b4-a455-4aed482fb906 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 770372eb-6b43-4e78-8d0e-72b8116ef617 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training InstructBLIP: Towards general-purpose vision-language models with instruction tuning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea0ef4b0-f621-4a27-99ac-9cade3a0fa49 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training The llama 3 herd of models.arXiv e-prints, pages arXiv–2407,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58521017-b486-4c5d-9016-065457d67803 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Taming transformers for high-resolution image synthesis, 2021
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44ff07d6-1c22-4cf0-9af8-e9d4234b7a7a · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Scaling rectified flow trans- formers for high-resolution image synthesis, 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38cf4303-dc7f-441a-8193-7b90a452f987 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training MME: A comprehensive evaluation benchmark for multimodal large language models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4db44839-f96d-4531-8039-1693ce4b100d · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Making llama see and draw with seed tokenizer, 2023
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8bcef66-4e17-4cc4-a754-1eb5699ce92a · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training When Attention Sink Emerges in Language Models: An Empirical View
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation f4be08c8-aa4a-428b-8a45-1e8138c42d1a · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Denoising diffu- sion probabilistic models, 2020
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 157e9bb7-b29d-4280-af17-4e1a4c143bc3 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Matryoshka query trans- former for large vision-language models, 2024
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed59ce8f-b34b-40a4-8569-5e2f1db43333 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Hudson and Christopher D
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afdfbab7-b92a-409a-9f15-ace553f9f1a4 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Unified language-vision pretraining in llm with dynamic discrete visual tokenization, 2024
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42b01c53-b1aa-48ea-879d-453a84b3f40c · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Seed-bench: Benchmarking multimodal llms with generative comprehension, 2023
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89ad5c31-298a-4fe0-8615-788fbc0fa72e · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Playground v2.5: Three insights towards enhancing aesthetic quality in text-to-image genera- tion, 2024
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98055ec9-bb71-4cd4-b80f-211c61d6739f · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Tokenpacker: Efficient visual projector for multimodal llm, 2024
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff3921b5-542c-46d0-9329-ce3915b223b9 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Evaluating object hallucination in large vision-language models, 2023
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57a49fa9-51e9-4bd3-beb3-457280b55024 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Llama-vid: An image is worth 2 tokens in large language models, 2023
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5eef721-4e82-4b5d-8470-f6b1b6d05811 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Boosting multimodal large language models with visual to- kens withdrawal for rapid inference, 2025
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb79f693-dbb4-4cf6-ab19-5298c61918be · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Improved baselines with visual instruction tuning, 2023
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c014e12-a6f6-4551-b5aa-80c8205add77 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Visual instruction tuning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 088db68e-ca28-4070-b332-7c9a29264fb8 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training World model on million-length video and language with blockwise ringattention
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c979e0a2-1a8d-47ba-90bc-5afa477b54bd · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Cheap and quick: Efficient vision- language instruction tuning for large language models, 2023
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8922ebe-c7cf-4373-a6e2-f50dcad78f6e · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Unitok: A uni- fied tokenizer for visual generation and understanding, 2025
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbe35b9d-02d7-4e02-a0e8-9838de9541ca · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation, 2024
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce0a3167-fc2c-44e5-aeae-5cdc2e55d129 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Learning transferable visual models from natural language supervision, 2021
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 591136e9-8ab0-4bb3-aecd-f36072565446 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Dynamicvit: Efficient vision transformers with dynamic token sparsification
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 785be7f9-7aa4-4659-93e8-89dcb2eaaf4a · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training High-resolution image syn- thesis with latent diffusion models, 2022
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7be0623-05e1-4c85-9800-da718381a4ed · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 37ab29a8-aa8c-4d0a-9d2b-2eea94331399 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Journeydb: A benchmark for generative im- age understanding.Advances in neural information process- ing systems, 36:49659–49678, 2023
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fddad9bf-4462-490c-a628-42401bd01eef · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Chameleon: Mixed-modal early-fusion foundation models, 2024
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5698df56-b8a2-4c2f-8be6-58ea4f5cf83e · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Visual autoregressive modeling: Scalable image generation via next-scale prediction, 2024
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 066be97a-8f72-4f14-bf20-9d31fc2136d2 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Llama: Open and efficient foundation lan- guage models, 2023
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98f9b6ff-325a-46dc-84d5-ef746b640c5d · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Neural discrete representation learning,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f816c86-b247-4de3-8ea8-20d3a015286c · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Emu3: Next-token prediction is all you need, 2024
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ab1316d-da58-49bc-8104-3f5b5c10c13e · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 510d55b9-c96c-44f8-8248-8422ac28ebf1 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Liquid: Language Models are Scalable and Unified Multi-modal Generators
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 6761f90c-f3e8-45aa-9e7c-f7abb418ae2c · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation af1e59da-4e95-4b36-a7f6-51e418c045df · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Efficient streaming language models with attention sinks.arXiv, 2023
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85ede03a-5ba7-4186-8ab4-3cffc56e2544 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 248ee663-a3b4-4695-83a8-cd9e1cb45e9e · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Scaling autoregressive multi-modal mod- els: Pretraining and instruction tuning, 2023
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bf6ecac-c1a9-4436-afa4-14e1ea516849 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Anygpt: Unified multimodal llm with discrete sequence modeling, 2025
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fcab27b-cfc4-4226-84c3-cffd89c4de0b · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training A-vl: Adaptive attention for large vision- language models, 2025
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3202d961-ff2e-4165-b8cb-f45f634bbcd5 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Llava-mini: Efficient image and video large multimodal models with one vision token, 2025
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24525677-03b9-475a-a718-cc04d98432cd · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Himix: Reducing computational com- plexity in large vision-language models, 2025
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce075547-819f-487b-a606-2e0bc49fa037 · outbound
On the Limits of Token Reduction for Efficient Unified Vision Language Training Ar- gus: A compact and versatile foundation model for vision
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.