Pith. sign in

Paper Citation Record · LEDGER

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer

As of 22 July 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2602.12286.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.12286 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T12:48:53.670343Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact13
  • verified fuzzy19
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cede7ccc-c9cd-4fb5-9ce2-25ec1532e5e8 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information pro- cessing systems, 35:23716–23736.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Flamingo: a visual language model for few-shot learning.Advances in neural information pro- cessing systems, 35:23716–23736

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:50:55.543957Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:adbd57ee5da9c12549b93587577de6d1a76dec058104f11bf1b895ecf07b02d6

Observation bc4966ec-5a15-456b-afd5-a32aa1e27991 · outbound

This paper cites X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:50:54.974623Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:0ead21df28c2b223bd05c1b657146ea89cfac5ba900a93b92ecb74bdd719bf50

Observation 030af5b8-e917-48e0-9908-63318f1f32d2 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual- linguistic tasks.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Internvl: Scaling up vision foundation models and aligning for generic visual- linguistic tasks

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:50:55.563985Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:b73f1fad10dd2be5f86c53ee9f2336453a5e5c242ef375654a452f84c8e9e183

Observation 5bd06aca-5da9-45c9-8ba0-09b7034d1941 · outbound

This paper cites Nucleotide transformer: building and evaluating robust foundation models for human genomics.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Nucleotide transformer: building and evaluating robust foundation models for human genomics

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:50:55.554839Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:e3a4b76bd3b927747b1d17b738b1210c7b20f93aaae991a56dee7a130c693cf5

Observation e32f8657-889e-4a48-8fe6-1670b1733959 · outbound

This paper cites A multimodal conversational agent for dna, rna and protein tasks.Nature Machine Intelligence, pages 1–14.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer A multimodal conversational agent for dna, rna and protein tasks.Nature Machine Intelligence, pages 1–14

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:50:55.541510Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:01b59837e1dd8a7029d49f18f744c9db598aa6e0bba4663c9260d678f9c3186a

Observation efd767d6-fae2-4828-8581-a401a3e03812 · outbound

This paper cites Genechat: A multi-modal large language model for gene function prediction.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Genechat: A multi-modal large language model for gene function prediction

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:50:55.536982Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:71686a733c860767ed2f4024a3c1377e0575a5d675f6fe37629d04f98d28bdca

Observation 3aa576ab-78f7-4dd9-9242-b9e0a8814869 · outbound

This paper cites Janusdna: A powerful bi-directional hybrid dna foundation model.arXiv preprint arXiv:2505.17257.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Janusdna: A powerful bi-directional hybrid dna foundation model.arXiv preprint arXiv:2505.17257

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:50:54.986933Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:f9caac1fb8a435bce0a88c8ad279157d9dfd762d22ca24fd66dc65094a8d2c86

Observation 074a0ee9-6299-4fbc-bfe6-ce8a59c5a8d5 · outbound

This paper cites Bioreason: Incentivizing multi- modal biological reasoning within a dna-llm model.arXiv preprint arXiv:2505.23579.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Bioreason: Incentivizing multi- modal biological reasoning within a dna-llm model.arXiv preprint arXiv:2505.23579

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:50:55.004072Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:f2c6a9916163b2ffef93e9329c69881ac31996574e523d54d13974bd7b64c27f

Observation 83ac525d-d89a-4dc5-ad46-42fdde1b8c14 · outbound

This paper cites Dnabert: pre-trained bidirectional en- coder representations from transformers model for dna- language in genome.Bioinformatics, 37(15):2112–2120.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Dnabert: pre-trained bidirectional en- coder representations from transformers model for dna- language in genome.Bioinformatics, 37(15):2112–2120

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:50:55.546053Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:deae142b8952333afa5cf348f70794ac24b905c7b8420acc22be9dd783ed9f95

Observation 88b86379-46da-429f-b762-734533c4d46b · outbound

This paper cites Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:50:55.539165Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:9a48f48e077126c855830b46b8894d0f80e4f41766768e4479179ea50eaeb133

Observation 91db1013-ff92-41f0-b6db-f489661d4d6c · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:50:55.571839Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:59c755b37a39b4b0b035c7db20addf298a86f5e7f777d75440f6ef51d082e3be

Observation 1cc5ccb8-20b8-4079-9d23-61ec287a0e8f · outbound

This paper cites Improved baselines with visual instruction tuning.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Improved baselines with visual instruction tuning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:50:55.534861Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:005c6b84e9c32c7614bcbc356afa1878c1bb2956e0e7a76a1ad4ab1db2e130d5

Observation 50c45b3b-825e-402f-98fe-6a7de3dd960d · outbound

This paper cites UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:50:54.993382Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:35988bb09818aa42529ca1c5e7279cbbcc1402a97c6538d6d8d2bf55a63cd773

Observation 992fbedf-4b7e-4fc3-bf17-5a091041fa30 · outbound

This paper cites A comprehensive overview of large language models.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer A comprehensive overview of large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:50:55.558836Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:f0aaab24d76038c565547c3d45df1c64f2758ef0e9bfe2e4cf05d39a940931a0

Observation 1e3e42e1-237d-410f-9276-92f32f005cd2 · outbound

This paper cites Generative ar- tificial intelligence for advancing discovery and design in biomateriomics.Intelligent Computing, 4:0117.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Generative ar- tificial intelligence for advancing discovery and design in biomateriomics.Intelligent Computing, 4:0117

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:50:55.548088Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:c8c039010ce7abe906bb03a3dccf45c299aa9fa738ba5042e686a1cbcc8593f6

Observation 63d18fe9-bd9a-4a27-a07c-edc9a37d5d2b · outbound

This paper cites Learning transferable visual models from nat- ural language supervision.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Learning transferable visual models from nat- ural language supervision

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:50:55.552698Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:bc4f2e690912f78782931916a43d4e753d2c938ab5717773f05f2a58e258f518

Observation 8f330759-e874-4fdc-beb7-1f063f459661 · outbound

This paper cites Context-aware regularization with markovian in- tegration for attention-based nucleotide analysis.arXiv preprint arXiv:2507.09378.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Context-aware regularization with markovian in- tegration for attention-based nucleotide analysis.arXiv preprint arXiv:2507.09378

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:50:54.975971Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:0f2880b5a0b3c13df70e7f440b21f5ba4983cfa2ee7f843002779cb5f94b3557

Observation 39a22a3c-849d-437f-a6c1-9e3054f90f36 · outbound

This paper cites Caduceus: Bi-directional equivariant long-range dna se- quence modeling.Proceedings of machine learning re- search, 235:43632.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Caduceus: Bi-directional equivariant long-range dna se- quence modeling.Proceedings of machine learning re- search, 235:43632

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:50:55.569516Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:bd1ff79508cdb3a957280e34a2d05173afb6d501b091d42a48116f4edabd35bf

Observation c7f39d7b-c367-466f-832d-a4b2584f7d91 · outbound

This paper cites Chatgpt: Optimizing language models for dialogue.OpenAI blog, 2(4).

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Chatgpt: Optimizing language models for dialogue.OpenAI blog, 2(4)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:50:55.556865Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:a42a43fe1dd8e69c490bb7514fedc0c47d3ac097300ca22f9a3285804662bd20

Observation 7c91389a-7dae-40df-a81f-9730bfca157e · outbound

This paper cites Neural machine translation of rare words with subword units.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Neural machine translation of rare words with subword units

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:50:55.550304Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:1eed4500c59cfb4815c7fd03d0bd18f14395aec7f986e545f8c7e7b7c728a4a7

Observation bba3e6f3-554a-478e-ba37-a5ca803aed9f · outbound

This paper cites PandaGPT: One Model To Instruction-Follow Them All.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer PandaGPT: One Model To Instruction-Follow Them All

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:50:54.983126Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:ffdd1081894b970f109285c27bda4893272911bf78c835d16a7003dac633bd50

Observation c6b5cc46-2869-4679-a03a-b10f9ad36b3d · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Emu: Generative Pretraining in Multimodality

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:22:11.462164Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:056f39bf59471ed21d7f62312d7fe3aa6bf725dc98fa23fb9fe5b45450d1c241

Observation 9f93216d-3c3a-494a-ba12-6a9063e84991 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:50:55.015355Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:f0825771b8dea74c1288d4b2666a11aa1386a8f949e54e1594d1e112cbc6d954

Observation d8db81ea-dea5-4161-8fbb-ad3b7326192e · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Attention is all you need.Advances in neural information processing systems, 30

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:50:55.574248Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:78e2e0e755c2df9ab98b6c42188d142cc8a44ff9beb8ce2a6c265a353a2c646e

Observation f47d4f76-1805-49e3-9c3f-295f98117fea · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Emu3: Next-Token Prediction is All You Need

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:50:54.990106Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:25c67fcf890c72a0850399892653d22b1feb2af1771f943ef6775e22b7e9fce5

Observation f59862fc-9abe-43b0-83f2-877f582d861a · outbound

This paper cites Omnireg-gpt: a high-efficiency foundation model for comprehensive genomic sequence understanding.Na- ture Communications, 16(1):10139.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Omnireg-gpt: a high-efficiency foundation model for comprehensive genomic sequence understanding.Na- ture Communications, 16(1):10139

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:50:55.576518Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:cd62089f2c6143bfbdbecae010b9b96f81727aa3628b3713092e60d55e8a51d3

Observation e031e524-0914-40b0-b861-658ab8f078c3 · outbound

This paper cites Genecom- pass: deciphering universal gene regulatory mecha- nisms with a knowledge-informed cross-species founda- tion model.Cell Research, 34(12):830–845.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Genecom- pass: deciphering universal gene regulatory mecha- nisms with a knowledge-informed cross-species founda- tion model.Cell Research, 34(12):830–845

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:50:55.561421Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:a2aabe588b1e56f7e81bef7858bb956851f4e8550a34a0769dd80cb3f2aa7384

Observation f03b4a70-2cbf-49ac-8a85-fb9e6a19f43e · outbound

This paper cites Qwen3 Technical Report.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Qwen3 Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:50:54.996941Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:1c11a69953cd5ddc810b302d5776490eae81f9866175e45860081601fdc88f95

Observation fb97339c-b57e-42ac-9e75-6dd7ce1648d2 · outbound

This paper cites Sigmoid loss for language image pre-training.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Sigmoid loss for language image pre-training

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:50:55.566781Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:2b78e33e9afa4789639274e1e83b03259e6a7cf3f320b13910472a6f1c3ac6d3

Observation 2a999c83-d40c-4fb3-b8c7-86768ed8ce66 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:50:55.000096Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:b901eb5d504e5d0de64ddcf1909b7357d06825dbd1a87795661853cd8efc784b

Observation ad21b9e0-6614-4731-9d1d-4c7867a4de31 · outbound

This paper cites PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:50:55.011715Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:bc68a2aafbf2b9d116471f258b4cf1f21443263191d67ce3e7cf1a748f8a18a0

Observation 6123afe3-c387-4df9-a7b7-d7a179826be7 · outbound

This paper cites Unified multimodal understanding and generation models: Advances, challenges, and opportunities.arXiv preprint arXiv:2505.02567.

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer Unified multimodal understanding and generation models: Advances, challenges, and opportunities.arXiv preprint arXiv:2505.02567

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:50:54.979634Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-05-16T12:48:53.670343Z digest=sha256:af9f857bd11835ac84613e3af11d202e884ef646174251db8b5c1848ddd003d8

Pith citing papers

No inbound Pith citation observations are available.