Pith. sign in

Paper Citation Record · LEDGER

On the Limits of Token Reduction for Efficient Unified Vision Language Training

As of 23 July 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2606.01503.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.01503 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T16:57:56.333595Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-20T06:30:07.809122+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bbfd7d18-a573-4d64-a687-e89dd65c2022 · outbound

This paper cites Token merging: Your vit but faster, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Token merging: Your vit but faster, 2023

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:50125c99f12179bb1f271c683fec4fe1513a142e20bc8c62002988437ae6ac83

Observation 5182b6d0-8fae-42f7-8347-08ec0a408930 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Sharegpt4v: Improving large multi-modal models with better captions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:fa6cf7a9bebd8a2f28402c853cedabf4695a0b23ba66bf3c43a538fac17e148a

Observation faa7894c-0022-48b4-a455-4aed482fb906 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024.

On the Limits of Token Reduction for Efficient Unified Vision Language Training An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:fc2a225a778cbdae1cf60d736c0fae52a897b11ca7a07ecbdee44f20039621ae

Observation 770372eb-6b43-4e78-8d0e-72b8116ef617 · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning.

On the Limits of Token Reduction for Efficient Unified Vision Language Training InstructBLIP: Towards general-purpose vision-language models with instruction tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:090b81e011a414ad13dd08d340a2a67b715905a3733cb16c4e53ebd5ba96384c

Observation ea0ef4b0-f621-4a27-99ac-9cade3a0fa49 · outbound

This paper cites The llama 3 herd of models.arXiv e-prints, pages arXiv–2407,.

On the Limits of Token Reduction for Efficient Unified Vision Language Training The llama 3 herd of models.arXiv e-prints, pages arXiv–2407,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:e1a113b860f52cc88688d964849af92a006cf1028897bf1af510832158983a27

Observation 58521017-b486-4c5d-9016-065457d67803 · outbound

This paper cites Taming transformers for high-resolution image synthesis, 2021.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Taming transformers for high-resolution image synthesis, 2021

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:c441a38a5473c2bccd3f9fe4daf4a05c4ea8d7f5e1afa79b19fad24ce06c95a8

Observation 44ff07d6-1c22-4cf0-9af8-e9d4234b7a7a · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis, 2024.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Scaling rectified flow trans- formers for high-resolution image synthesis, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:92384b8a1dc31e871340d599644201df53ded80c718996545f848f40a1390194

Observation 38cf4303-dc7f-441a-8193-7b90a452f987 · outbound

This paper cites MME: A comprehensive evaluation benchmark for multimodal large language models.

On the Limits of Token Reduction for Efficient Unified Vision Language Training MME: A comprehensive evaluation benchmark for multimodal large language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:e9f264167e7d208a685f833d66cae09273a811179a135f150b70b0d4d42cd65e

Observation 4db44839-f96d-4531-8039-1693ce4b100d · outbound

This paper cites Making llama see and draw with seed tokenizer, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Making llama see and draw with seed tokenizer, 2023

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:e2cdc2a217634d98da66b97b09b811b18387e427b66687a7f41a48d1fc88083a

Observation a8bcef66-4e17-4cc4-a754-1eb5699ce92a · outbound

This paper cites When Attention Sink Emerges in Language Models: An Empirical View.

On the Limits of Token Reduction for Efficient Unified Vision Language Training When Attention Sink Emerges in Language Models: An Empirical View

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:02:24.305996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:98e051cc32539b7d31fa654d52deeffc61db27d99ace383ce770da6ca151bebd

Observation f4be08c8-aa4a-428b-8a45-1e8138c42d1a · outbound

This paper cites Denoising diffu- sion probabilistic models, 2020.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Denoising diffu- sion probabilistic models, 2020

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:c9de18ca31835fa621ad878cbf3650739026768fbffd667edbc56ec050ed2a8b

Observation 157e9bb7-b29d-4280-af17-4e1a4c143bc3 · outbound

This paper cites Matryoshka query trans- former for large vision-language models, 2024.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Matryoshka query trans- former for large vision-language models, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:07013fea45dfb24dcdef9d811c92aeb12928586aaf204cadfba09e7392c62b28

Observation ed59ce8f-b34b-40a4-8569-5e2f1db43333 · outbound

This paper cites Hudson and Christopher D.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Hudson and Christopher D

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:9af9359f5035e1b5c90850abaa5b632442e36e359e532f21925e98a4375162ce

Observation afdfbab7-b92a-409a-9f15-ace553f9f1a4 · outbound

This paper cites Unified language-vision pretraining in llm with dynamic discrete visual tokenization, 2024.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Unified language-vision pretraining in llm with dynamic discrete visual tokenization, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:e030d2a9ce1f421b9a311c2c1e575dfa4a907481ed7aa35097bd24d9692702a6

Observation 42b01c53-b1aa-48ea-879d-453a84b3f40c · outbound

This paper cites Seed-bench: Benchmarking multimodal llms with generative comprehension, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Seed-bench: Benchmarking multimodal llms with generative comprehension, 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:dde2861843d461cc7a689bc1f25c3b9e31c0bd182275b752b929b8ade23bf011

Observation 89ad5c31-298a-4fe0-8615-788fbc0fa72e · outbound

This paper cites Playground v2.5: Three insights towards enhancing aesthetic quality in text-to-image genera- tion, 2024.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Playground v2.5: Three insights towards enhancing aesthetic quality in text-to-image genera- tion, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:ec49e53b6c21454395d8dca2f84dd6066a98188315b57f4d32485f458632d65e

Observation 98055ec9-bb71-4cd4-b80f-211c61d6739f · outbound

This paper cites Tokenpacker: Efficient visual projector for multimodal llm, 2024.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Tokenpacker: Efficient visual projector for multimodal llm, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:8a6910a32026fc13b7fe6975fbbea1450e75b1af3c20572f0fd6e53f3a35a17a

Observation ff3921b5-542c-46d0-9329-ce3915b223b9 · outbound

This paper cites Evaluating object hallucination in large vision-language models, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Evaluating object hallucination in large vision-language models, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:f5d769c35d61952b766fe3f8d52b970f1540e5bfa863b3798530ecbabb500dd2

Observation 57a49fa9-51e9-4bd3-beb3-457280b55024 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Llama-vid: An image is worth 2 tokens in large language models, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:5e48614ef10490e35191c7bf12275ae1e3e80cfeab1d08a4411344b117689e27

Observation d5eef721-4e82-4b5d-8470-f6b1b6d05811 · outbound

This paper cites Boosting multimodal large language models with visual to- kens withdrawal for rapid inference, 2025.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Boosting multimodal large language models with visual to- kens withdrawal for rapid inference, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:87bed4cefc38d60293dd3d5f28281d98cae719a1e6a47efffbeddb8933390d46

Observation fb79f693-dbb4-4cf6-ab19-5298c61918be · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Improved baselines with visual instruction tuning, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:5de91c9405a49ed4eb87fa1c7c7185e1251264cb78b58d3f661aa7a420a6b875

Observation 6c014e12-a6f6-4551-b5aa-80c8205add77 · outbound

This paper cites Visual instruction tuning.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Visual instruction tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:63298f2c20e08a0fe29e1ef609636b58f87fd7c719efe5e1f5ef8b3d23af0e00

Observation 088db68e-ca28-4070-b332-7c9a29264fb8 · outbound

This paper cites World model on million-length video and language with blockwise ringattention.

On the Limits of Token Reduction for Efficient Unified Vision Language Training World model on million-length video and language with blockwise ringattention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:44f04993473fd39aec45e3c31694c863905965b7512f03192e8df012ade93870

Observation c979e0a2-1a8d-47ba-90bc-5afa477b54bd · outbound

This paper cites Cheap and quick: Efficient vision- language instruction tuning for large language models, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Cheap and quick: Efficient vision- language instruction tuning for large language models, 2023

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:d600b95dabf99a39863653c8f8753694bb7fee08147e0199c5fa32bfa9f40279

Observation a8922ebe-c7cf-4373-a6e2-f50dcad78f6e · outbound

This paper cites Unitok: A uni- fied tokenizer for visual generation and understanding, 2025.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Unitok: A uni- fied tokenizer for visual generation and understanding, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:ca904445b9695d8b9082e7f468e8af05accdeae7e3c238b8eff0b3f5ea9eba6a

Observation bbe35b9d-02d7-4e02-a0e8-9838de9541ca · outbound

This paper cites Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation, 2024.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:aec1ae4a5923f22e06b4f1332936e7a7d834dfc9d428bcf7053d29e97fce28f1

Observation ce0a3167-fc2c-44e5-aeae-5cdc2e55d129 · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Learning transferable visual models from natural language supervision, 2021

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:4f3ca42061243e0eadaf6eaa6d04c00ca1446392cfbd0bd768569d3365b1efe2

Observation 591136e9-8ab0-4bb3-aecd-f36072565446 · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Dynamicvit: Efficient vision transformers with dynamic token sparsification

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:b3288d917d754447a96a92ae5e5319cf1bd726e0f509015b92b29f2840c5e4b4

Observation 785be7f9-7aa4-4659-93e8-89dcb2eaaf4a · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models, 2022.

On the Limits of Token Reduction for Efficient Unified Vision Language Training High-resolution image syn- thesis with latent diffusion models, 2022

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:40cadbf7198a8ae0af983b64c37aef02705e43e61dcda8e3a2d6d534fe640944

Observation e7be0623-05e1-4c85-9800-da718381a4ed · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:02:24.309040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:bebe07dc82d9f8ee0f36737b8d7827a17fe78549395416f1a9f99bf8cd80297f

Observation 37ab29a8-aa8c-4d0a-9d2b-2eea94331399 · outbound

This paper cites Journeydb: A benchmark for generative im- age understanding.Advances in neural information process- ing systems, 36:49659–49678, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Journeydb: A benchmark for generative im- age understanding.Advances in neural information process- ing systems, 36:49659–49678, 2023

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:1491ed463c9271db3e1e17714251856dacb840ddcb58bcde54df0311efacd2b5

Observation fddad9bf-4462-490c-a628-42401bd01eef · outbound

This paper cites Chameleon: Mixed-modal early-fusion foundation models, 2024.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Chameleon: Mixed-modal early-fusion foundation models, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:dbb4ffbec8795ba3d146d2fb3dafa6d1c2e0acdeee6f651da8c5d52df93fd15f

Observation 5698df56-b8a2-4c2f-8be6-58ea4f5cf83e · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction, 2024.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Visual autoregressive modeling: Scalable image generation via next-scale prediction, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:25122a04b0f6e30ecebd58180fbd8f1e38dee9730a53198100354fd25586c511

Observation 066be97a-8f72-4f14-bf20-9d31fc2136d2 · outbound

This paper cites Llama: Open and efficient foundation lan- guage models, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Llama: Open and efficient foundation lan- guage models, 2023

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:936572e07dde232b901215179d10a2c9921ad75d03d2cae41eae806ce227e7e7

Observation 98f9b6ff-325a-46dc-84d5-ef746b640c5d · outbound

This paper cites Neural discrete representation learning,.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Neural discrete representation learning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:15e046e1973c73d2ad5a3ebcc613f9a2ec1c1fac1309bc641b4028f9a6d71eb9

Observation 0f816c86-b247-4de3-8ea8-20d3a015286c · outbound

This paper cites Emu3: Next-token prediction is all you need, 2024.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Emu3: Next-token prediction is all you need, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:31d99834cab354be7e1e59c80b2237f533f83b8527204dc0c770ace111b28698

Observation 2ab1316d-da58-49bc-8104-3f5b5c10c13e · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:02:24.303829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:05db4fb512220a4176edd7cbd1a9ba618dcc2e04f896ec1fb657906d639a57b8

Observation 510d55b9-c96c-44f8-8248-8422ac28ebf1 · outbound

This paper cites Liquid: Language Models are Scalable and Unified Multi-modal Generators.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:02:24.295561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:778c5501be52cfca940939dbdac8534ad15b33baaba951133b41a24a11a9c7f9

Observation 6761f90c-f3e8-45aa-9e7c-f7abb418ae2c · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

On the Limits of Token Reduction for Efficient Unified Vision Language Training VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:02:24.298245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:7a8d8d9faf0f776fbd3282b7bc2ff4e5b837fb524f482c1933ae0f4bc9cb74e1

Observation af1e59da-4e95-4b36-a7f6-51e418c045df · outbound

This paper cites Efficient streaming language models with attention sinks.arXiv, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Efficient streaming language models with attention sinks.arXiv, 2023

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:6cc39cfc434f84f9ec94dd06d3025a54c374016ca6ac8fa208ff9cf2588ffd6c

Observation 85ede03a-5ba7-4186-8ab4-3cffc56e2544 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:02:24.300518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:75bd0c2113107507c02ce6e2b2b7f955675fab86e42ee0d4164590265794bec0

Observation 248ee663-a3b4-4695-83a8-cd9e1cb45e9e · outbound

This paper cites Scaling autoregressive multi-modal mod- els: Pretraining and instruction tuning, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Scaling autoregressive multi-modal mod- els: Pretraining and instruction tuning, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:df4bc14aacda1f4c5f7dd8e26e5e7dcb0983509e518b0f6f365305c02197d04c

Observation 5bf6ecac-c1a9-4436-afa4-14e1ea516849 · outbound

This paper cites Anygpt: Unified multimodal llm with discrete sequence modeling, 2025.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Anygpt: Unified multimodal llm with discrete sequence modeling, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:f95603eca31d3502bdf85013433ded95f39a05cf33af0a0ee657f41390a8aa87

Observation 3fcab27b-cfc4-4226-84c3-cffd89c4de0b · outbound

This paper cites A-vl: Adaptive attention for large vision- language models, 2025.

On the Limits of Token Reduction for Efficient Unified Vision Language Training A-vl: Adaptive attention for large vision- language models, 2025

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:9dfda97773be1e81bd971fffb51100f823e041a198e1ef3ac5d1f894b73ef5ee

Observation 3202d961-ff2e-4165-b8cb-f45f634bbcd5 · outbound

This paper cites Llava-mini: Efficient image and video large multimodal models with one vision token, 2025.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Llava-mini: Efficient image and video large multimodal models with one vision token, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:3d8c291fe81e8a0594375e4a3fae9bf210659578921a539217eacd2b555defb0

Observation 24525677-03b9-475a-a718-cc04d98432cd · outbound

This paper cites Himix: Reducing computational com- plexity in large vision-language models, 2025.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Himix: Reducing computational com- plexity in large vision-language models, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:ec7557ba8bbde7bf4f4c5be14159dd9c094a78f28068c9861dc3fb979c3f2d61

Observation ce075547-819f-487b-a606-2e0bc49fa037 · outbound

This paper cites Ar- gus: A compact and versatile foundation model for vision.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Ar- gus: A compact and versatile foundation model for vision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:c10f8dce2d5b1ccda4b5524886eb2ba856b57f04a8f52522603b33556fd86e10

Pith citing papers

No inbound Pith citation observations are available.