Pith. sign in

Paper Citation Record · LEDGER

Learning Transferable Visual Models From Natural Language Supervision

As of 23 July 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2103.00020.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2103.00020 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-20T06:30:07.809122+00:00

measured 100 of 318 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-15T13:27:51.848177Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T19:17:31.661936Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4611368c-8ba4-4019-b3a8-5a45525f647e · inbound

Diffusion Models Beat GANs on Image Synthesis cites this paper.

Diffusion Models Beat GANs on Image Synthesis Learning Transferable Visual Models From Natural Language Supervision

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-13T11:16:28.615424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T11:16:28.445702Z digest=sha256:eb829516a5d47f625d6eba8f65fe9706f045402cca64ec43d4d8f4e265b7ed27

Observation f081a67d-ec9e-435e-9352-d4eb81ec488b · inbound

LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs cites this paper.

LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs Learning Transferable Visual Models From Natural Language Supervision

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:21:01.083625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-12T10:21:01.062199Z digest=sha256:1303c66ebc6e631df243dac63251c7a8a21e52fef78c23a0a8d316a0759db7b2

Observation cc8e0c78-d932-4e80-8598-c0117f88e847 · inbound

Florence: A New Foundation Model for Computer Vision cites this paper.

Florence: A New Foundation Model for Computer Vision Learning Transferable Visual Models From Natural Language Supervision

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T09:38:09.547810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-16T09:38:09.427509Z digest=sha256:08def9a5c851447ed34d51d311a96d69061c68217b3dc59914f056c6db902a4b

Observation f2b65885-e095-4b6d-85ba-6a6373f07964 · inbound

GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models cites this paper.

GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models Learning Transferable Visual Models From Natural Language Supervision

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:56:11.433509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-11T05:56:10.970591Z digest=sha256:e6bcb2af151b5b3e888c9ddb586c879aae7519e6daf6b3ccddb59ef27836d53a

Observation ea90e943-4f33-4574-a42b-2cf64b4103d5 · inbound

Text and Code Embeddings by Contrastive Pre-Training cites this paper.

Text and Code Embeddings by Contrastive Pre-Training Learning Transferable Visual Models From Natural Language Supervision

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T19:24:12.033622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-15T19:24:11.907204Z digest=sha256:da87dad622f326f7d086c3bb11d13b6dbff5815f12672a2087065a484d0b76f4

Observation 5236b9e0-2f60-4fd7-a69a-ebb033bde570 · inbound

Hierarchical Text-Conditional Image Generation with CLIP Latents cites this paper.

Hierarchical Text-Conditional Image Generation with CLIP Latents Learning Transferable Visual Models From Natural Language Supervision

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-10T16:55:57.859246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-10T16:55:57.612364Z digest=sha256:c0839ca7451f1b024c6b39625e0f71bf4292bde26bee3ee2dcf445abb06a8c6a

Observation c3bc78d9-7dc0-4754-be2b-33cedf04a679 · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning Learning Transferable Visual Models From Natural Language Supervision

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:22:30.554692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:593132abda3f75ed0068347798fec1b99e76210f4416ee4cbb82401a42f36fe4

Observation af6e434d-dc93-4565-88a0-9aab6ca65ac4 · inbound

Scaling Laws and Interpretability of Learning from Repeated Data cites this paper.

Scaling Laws and Interpretability of Learning from Repeated Data Learning Transferable Visual Models From Natural Language Supervision

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T15:52:40.444984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=arxiv_source observed=2026-05-17T15:52:40.335080Z digest=sha256:5db30457a4b431924667a1719c16124887a47aec8a30d8021f306a68d4017d30

Observation 25666593-e016-4a69-b999-34abf94ae959 · inbound

An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion cites this paper.

An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion Learning Transferable Visual Models From Natural Language Supervision

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:08:55.619148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-11T18:08:55.311069Z digest=sha256:eb7204dd5182d3d66e888478a9e0e02b0b1bfd942703c1ee97a6967ceeccb088

Observation 007d1e49-c192-4450-8dbd-d6745f5f1965 · inbound

Prompt-to-Prompt Image Editing with Cross Attention Control cites this paper.

Prompt-to-Prompt Image Editing with Cross Attention Control Learning Transferable Visual Models From Natural Language Supervision

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:00:01.546299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-11T07:00:01.154743Z digest=sha256:7077bbf92631566df41b69404b1e283bdb49c0406fed145c592112c6b83ff66b

Observation 8cc47183-97ee-4b2a-ad1f-28f1224ada5c · inbound

DetailCLIP: Injecting Image Details into CLIP's Feature Space cites this paper.

DetailCLIP: Injecting Image Details into CLIP's Feature Space Learning Transferable Visual Models From Natural Language Supervision

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-24T11:09:22.331991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-24T11:08:20.298043Z digest=sha256:07165929887c99ca6e0431700b1e163f9db71d48b1c96a3f5747dbcf1883fef9

Observation ca55c9cb-afb3-4525-b1ab-0b31d8976616 · inbound

LAION-5B: An open large-scale dataset for training next generation image-text models cites this paper.

LAION-5B: An open large-scale dataset for training next generation image-text models Learning Transferable Visual Models From Natural Language Supervision

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-13T14:22:17.364966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T14:22:16.968028Z digest=sha256:f5543df47abd3030c3b254212031722342b08490cdca62e251490d9e3b4863bc

Observation 23f28d1d-01fb-461c-a5a0-e75cad7cb7dc · inbound

Editing Models with Task Arithmetic cites this paper.

Editing Models with Task Arithmetic Learning Transferable Visual Models From Natural Language Supervision

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-05-13T08:09:13.164829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T08:09:12.716163Z digest=sha256:47ea98da8e0e9f5509eafbe9bfdc21f63371f163414cc7324e34e7e5135df9f0

Observation 34fe5f8a-f436-4989-bc98-c56c48ae7a9f · inbound

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models cites this paper.

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models Learning Transferable Visual Models From Natural Language Supervision

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T00:10:49.712361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-12T00:10:48.610351Z digest=sha256:66db1a8e195fb3d29438ee93321781c7dc685fa385ac5373469d46a7a49f0c8b

Observation a3f9ffef-f155-42b6-a10f-88863cba22bf · inbound

Visual Instruction Tuning cites this paper.

Visual Instruction Tuning Learning Transferable Visual Models From Natural Language Supervision

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:22:03.479963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-11T08:22:03.403362Z digest=sha256:5149402fc63d43e470b8911413ba8e8a3807508dab10f7e734bbe669a71fa834

Observation 9dd489dc-b93c-4343-945c-41bfa9d82804 · inbound

Shap-E: Generating Conditional 3D Implicit Functions cites this paper.

Shap-E: Generating Conditional 3D Implicit Functions Learning Transferable Visual Models From Natural Language Supervision

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:32:06.844682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-16T15:32:06.563955Z digest=sha256:1c2c3ac40029d925d28be0845e7b2a0fb65b9942fc922325ecefec31c0d2d35e

Observation 3580bb70-9748-442a-b2cd-b5a1e6346a01 · inbound

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models cites this paper.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Learning Transferable Visual Models From Natural Language Supervision

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T09:55:35.527784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:6c6f165fe8f351774b8cbdd90293e5c07e843512c933ffa485cb433fab0f87e4

Observation c858f35a-616e-4ac5-a1ab-b8929954dd46 · inbound

Training Diffusion Models with Reinforcement Learning cites this paper.

Training Diffusion Models with Reinforcement Learning Learning Transferable Visual Models From Natural Language Supervision

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:16:31.082821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-11T20:16:30.840184Z digest=sha256:2dcd0aa91901b55c04fa39a5db783cb1baa4c9cf4d5244f189471d8d812223f3

Observation 65e679f7-33f3-4619-8979-8cac7b2e2bb2 · inbound

LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day cites this paper.

LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day Learning Transferable Visual Models From Natural Language Supervision

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-24T08:34:11.763176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-24T08:33:16.436778Z digest=sha256:8c013b15d7654bd09024893e43fd8973a5a102246ea07646b91e3c7f64033f91

Observation cf8d8144-a127-4595-8789-ad8317597b3a · inbound

SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis cites this paper.

SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis Learning Transferable Visual Models From Natural Language Supervision

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-10T15:22:03.580094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-10T15:22:03.530707Z digest=sha256:de38565f4f2c44fa700297f472c18c4acd2d9b79fc52d09857ddaf5f2c5c5b10

Observation fe608c0e-37f4-402c-b18a-61972b72cc02 · inbound

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution cites this paper.

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution Learning Transferable Visual Models From Natural Language Supervision

Reference 266

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T08:12:31.145802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=arxiv_source observed=2026-05-16T08:12:30.984870Z digest=sha256:ec55e1610a26c6d870190520e07eec3e3e36b97aff17720ccbf3e9ecdebd31d3

Observation 13c17552-54f6-47fd-b563-521e5f36b821 · inbound

Improved Baselines with Visual Instruction Tuning cites this paper.

Improved Baselines with Visual Instruction Tuning Learning Transferable Visual Models From Natural Language Supervision

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-12T19:11:33.940441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-12T19:11:33.783746Z digest=sha256:ab56a126d6ce842705f3aff14ee5d08c8b059b9e4cf775be4d498cd023d54c08

Observation 7d7aaa5e-a0d0-480e-9ad9-22ff8f14c19e · inbound

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets cites this paper.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Learning Transferable Visual Models From Natural Language Supervision

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:52.074775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:41e81ed34f1213970d414c49f3446b4d67040ba11e2ac9a95ce219bf947c5f71

Observation 1a8dadfc-ac93-450d-bcc9-dae5c939b830 · inbound

Massive Activations in Large Language Models cites this paper.

Massive Activations in Large Language Models Learning Transferable Visual Models From Natural Language Supervision

Reference 144

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:54.038805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:c2c0ff0f4ae50ba65e9c7c66160ea8a86c2e88eb99b0fad1c8da8691dd11a57c

Observation 2311dd82-29ef-42dd-bdec-49f171929ff3 · inbound

DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset cites this paper.

DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset Learning Transferable Visual Models From Natural Language Supervision

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:51:18.622178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-11T05:51:18.508352Z digest=sha256:9e554ab0c2cb30c8e4262d6687bca791612308d158bf9aad3ffc7bf8d2945010

Observation 8cbce99b-6f5a-4c61-8a46-bae82c5c5f87 · inbound

ANCHOR: LLM-driven Subject Conditioning for Text-to-Image Synthesis cites this paper.

ANCHOR: LLM-driven Subject Conditioning for Text-to-Image Synthesis Learning Transferable Visual Models From Natural Language Supervision

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-24T01:43:42.955185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-24T01:42:10.041654Z digest=sha256:d58b28c96a3274fdec37286d10683e4c18f3479fc32e51d633853ff73e2e396a

Observation 86f94a56-3e03-459d-a0d2-0509809a117a · inbound

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model cites this paper.

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model Learning Transferable Visual Models From Natural Language Supervision

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:57:26.970729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T05:57:26.887069Z digest=sha256:b2f78b1699b5ffd49705cbab5ef2a1441109d29e2cfe07b2c82590aa302c246d

Observation 6c6be21a-39d7-499a-b183-607ee0821150 · inbound

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models cites this paper.

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models Learning Transferable Visual Models From Natural Language Supervision

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-18T02:48:45.122729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T02:48:44.900467Z digest=sha256:539aece31ec86f52f6e747bddd8dc810b28bbc0ca9acf7ba73711f5862657f11

Observation 36efdace-0de0-45f0-98de-67ff2ec08aa2 · inbound

RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields cites this paper.

RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields Learning Transferable Visual Models From Natural Language Supervision

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:52:43.897262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-23T07:49:37.878395Z digest=sha256:77b66eadb7a138667eb998d9bd752dc77538652938c34f5caea1488bdf2fb83e

Observation 0536ba39-0e97-458c-bcaa-3a12771ed14a · inbound

Multimodal Contextualized Support for Enhancing Video Retrieval System cites this paper.

Multimodal Contextualized Support for Enhancing Video Retrieval System Learning Transferable Visual Models From Natural Language Supervision

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:15:28.473374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-23T07:14:34.843867Z digest=sha256:9a5119538ee64cf7c16a116535cea012d4e5628f347b80d4f454e502e508fc95

Observation dd07465c-8fc5-4f67-9c60-9add8f23d751 · inbound

3D Foundation Model for Generalizable Disease Detection in Head Computed Tomography cites this paper.

3D Foundation Model for Generalizable Disease Detection in Head Computed Tomography Learning Transferable Visual Models From Natural Language Supervision

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T03:27:26.841683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:be3a07a627a72b70149389b102e9a05bc657fddaa63e9eb699e6cb3914240a8b

Observation 89c5a0e1-1b85-480b-9024-db7f29dcf793 · inbound

Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models cites this paper.

Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models Learning Transferable Visual Models From Natural Language Supervision

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T02:55:19.772598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-23T02:53:34.960236Z digest=sha256:15a8bf9b6ad846560c8faece00a1f9e77b184e0d79d4b2193dc222051c2b2b83

Observation 0d6564e6-fd8b-4463-8173-6551cbcc1ad0 · inbound

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data cites this paper.

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data Learning Transferable Visual Models From Natural Language Supervision

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:55:52.193763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-17T20:55:52.109166Z digest=sha256:369e9c8c47e1f5c82813658069b8b35520a49eb19ffcd4cac728b5d12a8d6238

Observation 5ddf48c6-d170-4098-b627-d546fb0dba2d · inbound

v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning cites this paper.

v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning Learning Transferable Visual Models From Natural Language Supervision

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:37:17.465576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-19T12:36:32.030301Z digest=sha256:0351c842a689d1f19bb5d2cb9735468b4f31df23e64ff29259880f2a03963ab7

Observation f8df1256-cd55-4e25-9528-5148b6c97d21 · inbound

Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration cites this paper.

Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration Learning Transferable Visual Models From Natural Language Supervision

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T12:52:18.057828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-19T12:48:44.324236Z digest=sha256:5f509aa3b87b365ca038a2779b9f4bb46f9ba4ea2610b9a8840e8b13b2d9065f

Observation 9835badb-423e-4611-be3a-170d6b1fff49 · inbound

Structured Labeling Enables Faster Vision-Language Models for End-to-End Autonomous Driving cites this paper.

Structured Labeling Enables Faster Vision-Language Models for End-to-End Autonomous Driving Learning Transferable Visual Models From Natural Language Supervision

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:20:50.505876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-22T00:16:35.823270Z digest=sha256:0480fffbb8279e9444c21190bcd23eda947529f6854c127f44edccd6ea8d82b6

Observation 00ac0409-1e37-4d5d-a77b-1c419f602110 · inbound

HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding cites this paper.

HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding Learning Transferable Visual Models From Natural Language Supervision

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-19T10:47:15.070458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-19T10:44:01.880405Z digest=sha256:5bc1b0fc790f264c31aeb9709e76a3cf458ad756ea3fbd4def5ba00f96986b77

Observation a7a59f63-7e06-4afe-b57d-b4a09d21eb2d · inbound

Self-Predictive Representations for Combinatorial Generalization in Behavioral Cloning cites this paper.

Self-Predictive Representations for Combinatorial Generalization in Behavioral Cloning Learning Transferable Visual Models From Natural Language Supervision

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:17:14.168233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-19T09:15:45.511104Z digest=sha256:d538fbea780fdafe35b40af252c4ddc3dab6ee718141394a60d0a5c6420d22dc

Observation 22bfa6a9-424c-4e61-8d81-4f05bb776205 · inbound

CLIP the Landscape: Automated Tagging of Crowdsourced Landscape Images cites this paper.

CLIP the Landscape: Automated Tagging of Crowdsourced Landscape Images Learning Transferable Visual Models From Natural Language Supervision

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T09:03:02.191639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-19T09:02:18.576084Z digest=sha256:e0c10a1a1c5b41b3b8f03e027454412e0d36e9619c2d366b51542889aaab993a

Observation 5d2ea360-17d8-46bd-9cfa-05e089cab8b2 · inbound

Navigating the Challenges of AI-Generated Image Detection in the Wild: What Truly Matters? cites this paper.

Navigating the Challenges of AI-Generated Image Detection in the Wild: What Truly Matters? Learning Transferable Visual Models From Natural Language Supervision

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-21T23:34:26.585924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-21T23:31:40.691896Z digest=sha256:0d4ddb7f3f1d16cb8ac655b93de208f454a4db1ce22649fecd730505f8851a1b

Observation 13af1886-83a2-42b6-a1f6-e5e64f0d3453 · inbound

Scalable Option Learning in High-Throughput Environments cites this paper.

Scalable Option Learning in High-Throughput Environments Learning Transferable Visual Models From Natural Language Supervision

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:06:49.685289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=arxiv_source observed=2026-05-18T20:04:58.064472Z digest=sha256:fa24864651784253871ac769dd859407ff559ace73c82292408111fb7a4ec5f5

Observation 1e1dd5c8-ef24-4e4c-bd76-d30d8ea2f5a8 · inbound

Synthetic Homes: A Multimodal Generative AI Pipeline for Residential Building Data Generation under Data Scarcity cites this paper.

Synthetic Homes: A Multimodal Generative AI Pipeline for Residential Building Data Generation under Data Scarcity Learning Transferable Visual Models From Natural Language Supervision

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T17:11:39.036438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T17:10:57.875842Z digest=sha256:17dcf01981b21151e85f5c541aa64fcbca3eb348c0d45dd5f12551c21ede0499

Observation 9fe78699-19dc-4517-90de-6e058705e9e6 · inbound

VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis cites this paper.

VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis Learning Transferable Visual Models From Natural Language Supervision

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-18T15:31:33.267640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T15:30:29.933558Z digest=sha256:b5d8580ad130c5e3303eed3a47213df801d67441cfc170203b75befc0b72311d

Observation 573a938d-86b4-4ae1-9665-b7fb893383ff · inbound

Artificial Phantasia: Emergent Mental Imagery in Large Language Models cites this paper.

Artificial Phantasia: Emergent Mental Imagery in Large Language Models Learning Transferable Visual Models From Natural Language Supervision

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:44:22.810494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=arxiv_source observed=2026-05-21T21:41:39.111769Z digest=sha256:ca690c8cf3dd2d817ed3d2f6fd1850e2b2946ceddbc30d078bb4a7454daf81b8

Observation 7a8de36d-cf3c-47c7-b551-0198db56fbfc · inbound

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance cites this paper.

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance Learning Transferable Visual Models From Natural Language Supervision

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:23.778353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=arxiv_source observed=2026-05-18T13:10:18.700497Z digest=sha256:ab71be93188187f47355925c51d99dd0add5a6aabc57ee5c1307e507e8e65ecb

Observation 2ebfdd4c-ef55-4d33-b2cf-7a3196c0f830 · inbound

SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning cites this paper.

SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning Learning Transferable Visual Models From Natural Language Supervision

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-21T20:24:21.299776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-21T20:24:02.748854Z digest=sha256:2aab454a84bc22fa3f92316dc56aece775161e9142d4afcb2aeec307cc2716dd

Observation 63c99a37-76a1-44e3-bc1c-42715cf8e5d3 · inbound

Foundation Models for Discovery and Exploration in Chemical Space cites this paper.

Foundation Models for Discovery and Exploration in Chemical Space Learning Transferable Visual Models From Natural Language Supervision

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:52:24.969807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:52:10.848118Z digest=sha256:227a25a27bdc8936252e4c22662095a991708fc208d1403527584677cd84d203

Observation c530b18e-5361-44f7-8dd4-d491bb708b70 · inbound

MM-Telco: Benchmarks and Multimodal Large Language Models for Telecom Applications cites this paper.

MM-Telco: Benchmarks and Multimodal Large Language Models for Telecom Applications Learning Transferable Visual Models From Natural Language Supervision

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:10:23.003051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-17T22:06:30.391838Z digest=sha256:7e7faca7830ba138261febd402567f1e202329d652055aa73bf32cb0dda1290c

Observation 12b949af-635b-41ff-8ded-891fdfbb580c · inbound

Physics-Based Benchmarking Metrics for Multimodal Synthetic Images cites this paper.

Physics-Based Benchmarking Metrics for Multimodal Synthetic Images Learning Transferable Visual Models From Natural Language Supervision

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T21:15:16.419195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-17T21:12:38.608993Z digest=sha256:56f556087144fba0e70e152eace0b732f1c12e37497c5da84b7626f0eff6a2b4

Observation 0f33d222-d042-49f0-afd6-eb99ec39644d · inbound

Developing an AI Course for Synthetic Chemistry Students cites this paper.

Developing an AI Course for Synthetic Chemistry Students Learning Transferable Visual Models From Natural Language Supervision

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T06:14:09.568297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:fc51338a1bd837553e6e87fc998b0b2674c4753190951c4c443d1b9599f78698

Observation d59da5a5-6d7b-416e-b2a1-8ef97cc6cef9 · inbound

Optical Context Compression Is Just (Bad) Autoencoding cites this paper.

Optical Context Compression Is Just (Bad) Autoencoding Learning Transferable Visual Models From Natural Language Supervision

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T02:58:55.156339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-17T02:55:06.012687Z digest=sha256:20c87553ca1c466ce2dfab095076e5938bddfc7c5b1adc8faf472154ee50c9e8

Observation 49b20c33-1a91-4200-b12f-f634c16899bd · inbound

A Simple Method to Enhance Pre-trained Language Models with Speech Tokens for Classification cites this paper.

A Simple Method to Enhance Pre-trained Language Models with Speech Tokens for Classification Learning Transferable Visual Models From Natural Language Supervision

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:48:45.954709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=arxiv_source observed=2026-05-17T00:46:08.921194Z digest=sha256:4c9ebcdf7ca17bc52a3ae326528564b6acdf36432a77ae699a0afacbdf127964

Observation c69f4753-5f9a-403c-ae72-7ca38cd82d90 · inbound

Are vision-language models ready to zero-shot replace supervised classification models in agriculture? cites this paper.

Are vision-language models ready to zero-shot replace supervised classification models in agriculture? Learning Transferable Visual Models From Natural Language Supervision

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T21:18:32.270046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=arxiv_source observed=2026-05-16T21:15:20.717705Z digest=sha256:711c509af16d1cb0545682b90f252f6e29e9e389adf2031c812a33ceeab83af0

Observation e4a3408f-84b0-469a-978b-a02826524f9a · inbound

Flexible Multitask Learning with Factorized Diffusion Policy cites this paper.

Flexible Multitask Learning with Factorized Diffusion Policy Learning Transferable Visual Models From Natural Language Supervision

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:23:19.675360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-16T19:21:52.191786Z digest=sha256:8753896af19cc8a2537b0bdd2d779b5573ba38d29e18bc1ae0589d3be80898c5

Observation 4c1446c7-8ab3-4ad7-87c3-3743ea2b07df · inbound

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation cites this paper.

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation Learning Transferable Visual Models From Natural Language Supervision

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T19:48:21.903351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-16T19:43:37.604351Z digest=sha256:8ba8c4496d92c565139519e4baba15c2bbe87c0a93bfbfcecec50af8e133bd80

Observation e249cf04-29bc-45c9-b7f9-e9e4095fd6d5 · inbound

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation cites this paper.

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation Learning Transferable Visual Models From Natural Language Supervision

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T16:14:15.278212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-21T16:10:31.015783Z digest=sha256:ca6fc2f93ec34281dfa88045af49e225bbeb9a30768ea4dcf61e0af9a87cae74

Observation 6e650150-7356-4035-bec8-2dc9eb94747d · inbound

ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch cites this paper.

ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch Learning Transferable Visual Models From Natural Language Supervision

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-16T13:12:54.883047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-16T13:12:01.889341Z digest=sha256:5e2fc4ec65a09a804ba94ccc747b4d0f839c49f47eda4e9a7592816350dc9b1a

Observation a9ada99d-8c1a-4df9-b57f-c62c0e8b9ed6 · inbound

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding cites this paper.

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding Learning Transferable Visual Models From Natural Language Supervision

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T10:17:44.651201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-16T10:14:15.589472Z digest=sha256:05a2488a6059f0921cfc17e7ad578d1e3b9e919845970cf7c466d0496d5b2069

Observation 196c736b-ead6-4e52-ae12-b3a39e837c5b · inbound

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models cites this paper.

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models Learning Transferable Visual Models From Natural Language Supervision

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T14:50:14.771137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-21T14:48:21.787919Z digest=sha256:cdb102ee00b91c46aff31878d29a22d28ba4b5d59bb3d3f97a130025c36210c1

Observation 5593d480-f91a-4828-a670-cdbd195448f3 · inbound

Mitigating Long-Tail Bias via Prompt-Controlled Diffusion Augmentation cites this paper.

Mitigating Long-Tail Bias via Prompt-Controlled Diffusion Augmentation Learning Transferable Visual Models From Natural Language Supervision

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:30:44.322510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-16T07:28:46.372389Z digest=sha256:9a35494402c8b0f3aa005c4dc697df694ac7349064cc5370ea8e29f5237a6955

Observation a657a264-ee80-4195-9317-8f45d7c7a055 · inbound

HLGFA: High-Low Resolution Guided Feature Alignment for Unsupervised Anomaly Detection cites this paper.

HLGFA: High-Low Resolution Guided Feature Alignment for Unsupervised Anomaly Detection Learning Transferable Visual Models From Natural Language Supervision

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:07:11.864406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-16T03:05:37.320403Z digest=sha256:f98d5b9e86d44a1efb25ee38682c9b2bc1758163233d369ffb261ee5e1bd86ab

Observation 8a8bfa48-8018-43b3-a229-010dc4db4df4 · inbound

VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction cites this paper.

VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction Learning Transferable Visual Models From Natural Language Supervision

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-22T11:11:27.491007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-22T11:06:38.102348Z digest=sha256:aa6fc936ac25f03024d9522743e8f261a158a788513d0841aec2314b79fb9f28

Observation 562ada84-e78e-4096-a8d4-461c63d5d29c · inbound

InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation cites this paper.

InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation Learning Transferable Visual Models From Natural Language Supervision

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:10:15.871401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-15T19:07:46.733450Z digest=sha256:301083bcc56424c215bcdcfac49b40db6e536aaf3249683f56be78847e6bfdd5

Observation 0a67e26c-672a-4ee7-b554-436f09bc0646 · inbound

TiledAttention: a CUDA Tile SDPA Kernel for PyTorch cites this paper.

TiledAttention: a CUDA Tile SDPA Kernel for PyTorch Learning Transferable Visual Models From Natural Language Supervision

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T17:46:23.691773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-15T17:44:34.166709Z digest=sha256:2e61416176209399f5afeff551f50eb9c59aabf969e3317e5867a690cd258df9

Observation ea923910-c5aa-429e-b837-a0bcc5fba91a · inbound

Expectation and Acoustic Neural Network Representations Enhance Music Identification from Brain Activity cites this paper.

Expectation and Acoustic Neural Network Representations Enhance Music Identification from Brain Activity Learning Transferable Visual Models From Natural Language Supervision

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-21T11:40:03.503740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-21T11:36:06.967549Z digest=sha256:ca808e1b75241e348685dbf9ab4cc78ee76fd925afd502e61a1e313a9c925c6f

Observation 60d88704-1585-471e-80e3-6b71eb435ab5 · inbound

Vision Language Models Cannot Reason About Physical Transformation cites this paper.

Vision Language Models Cannot Reason About Physical Transformation Learning Transferable Visual Models From Natural Language Supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-15T13:27:51.848177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:27:51.848177Z digest=sha256:75bc345ee346091a4ef4110a38ec568696addde255c5a7f7ae9d1317473984c3

Observation 62e6e2c2-d18c-4f86-bfa5-17782c4b864a · inbound

TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization cites this paper.

TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization Learning Transferable Visual Models From Natural Language Supervision

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:16:09.742572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-15T15:12:23.459579Z digest=sha256:000dc329000bcc1b4e977bc8147d1f061825f64e216f4b7812f317e519797e6d

Observation 6326fe18-38ee-4ace-b3b5-8ff12724286e · inbound

Beyond Relevance: On the Relationship Between Retrieval and RAG Information Coverage cites this paper.

Beyond Relevance: On the Relationship Between Retrieval and RAG Information Coverage Learning Transferable Visual Models From Natural Language Supervision

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-15T13:10:49.476223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-15T13:10:06.367994Z digest=sha256:71d17a0d2996ed0e8176bc9784a1e0e5be018cdac5ce6b9788242ecaa80f3fed

Observation da59b9ee-14fe-4e1a-863e-68a40f065695 · inbound

Causal Attribution via Activation Patching cites this paper.

Causal Attribution via Activation Patching Learning Transferable Visual Models From Natural Language Supervision

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-21T11:10:02.094043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-21T11:09:50.326389Z digest=sha256:4f87f40f3ac6d22ce17b00997ad4affe1e90442ef8163fe06c3b6d2fe7c3dd11

Observation 0dffadd4-9b77-4b72-8501-6808e2bc6240 · inbound

Setting-Matched and Semantics-Scaled Benchmarking of One-Step Generative Models Against Multistep Diffusion and Flow Models cites this paper.

Setting-Matched and Semantics-Scaled Benchmarking of One-Step Generative Models Against Multistep Diffusion and Flow Models Learning Transferable Visual Models From Natural Language Supervision

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T12:20:00.134325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-15T12:15:38.186914Z digest=sha256:e874ab47c4365370a527527e953b2d27d5900e8d60b37c98dfd075aeb81bbecd

Observation 124979e9-15b7-4c01-b580-d78015c6331b · inbound

XNote: Benchmarking Automated Community Notes Generation for Image-based Contextual Deception cites this paper.

XNote: Benchmarking Automated Community Notes Generation for Image-based Contextual Deception Learning Transferable Visual Models From Natural Language Supervision

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-21T10:40:00.521100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-21T10:39:09.209094Z digest=sha256:7d77cf882f1a711008da95d279f6f41133a233dba235447ea352d3e3c327fedb

Observation e44e5e6e-1d4b-43ef-a85c-17f7509408e1 · inbound

The Gait Signature of Frailty: Transfer Learning based Deep Gait Models for Scalable Frailty Assessment cites this paper.

The Gait Signature of Frailty: Transfer Learning based Deep Gait Models for Scalable Frailty Assessment Learning Transferable Visual Models From Natural Language Supervision

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T00:23:22.785721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-15T00:21:21.413771Z digest=sha256:414c264d87006f726713cfb199f1a69e281ae4aa841ea2316f9392be5babc7c2

Observation 62f64836-03a6-4b6d-83cb-77e6c447b09f · inbound

Perceptual misalignment of texture representations in convolutional neural networks cites this paper.

Perceptual misalignment of texture representations in convolutional neural networks Learning Transferable Visual Models From Natural Language Supervision

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-21T09:29:57.028965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-21T09:27:00.103225Z digest=sha256:a5c11e0605473e712e3faaad601ce8ffeea8b0aeef22c013051c918a8a953213

Observation dd57b7d5-413d-4890-88a9-85f7d154b0e9 · inbound

Woosh: A Sound Effects Foundation Model cites this paper.

Woosh: A Sound Effects Foundation Model Learning Transferable Visual Models From Natural Language Supervision

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:53:15.829257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T20:51:08.144573Z digest=sha256:a270c72b4dae4b6837e8b3c2719b29ee54ce4fe71ffa2724e85806ffbcadd674

Observation d45f7f3a-a8f1-4df0-bf55-89fee9f99767 · inbound

Self-Directed Task Identification cites this paper.

Self-Directed Task Identification Learning Transferable Visual Models From Natural Language Supervision

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:43:19.028893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T21:40:16.017679Z digest=sha256:f0b974b68984009305725994d295c722ed5f1b16d318ee1ff3b5901ca1ee5e12

Observation 3f9113ee-c73b-4795-a664-c6f733e320d5 · inbound

An Explainable Vision-Language Model Framework with Adaptive PID-Tversky Loss for Lumbar Spinal Stenosis Diagnosis cites this paper.

An Explainable Vision-Language Model Framework with Adaptive PID-Tversky Loss for Lumbar Spinal Stenosis Diagnosis Learning Transferable Visual Models From Natural Language Supervision

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:43:18.798440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T21:43:13.507310Z digest=sha256:cad4825313c00c6f2c638b16d71bcc8b75e1d3976271212a57803432feacd57e

Observation 690d259a-8695-4b4b-963d-436a431aab69 · inbound

DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection cites this paper.

DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection Learning Transferable Visual Models From Natural Language Supervision

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:33:16.772776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T20:30:16.163179Z digest=sha256:9300d8f2b29786eedee1291ff67346fc41ae1411d149391df19d849b2f561d41

Observation e2d28930-f689-4c96-8961-957ba2c9946a · inbound

DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection cites this paper.

DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection Learning Transferable Visual Models From Natural Language Supervision

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:49:29.319724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-14T21:48:49.892964Z digest=sha256:0a11388f4ef5676643199379592196276cff7b3dfd4ef63e03b1e90530c87e48

Observation c8361e9c-74db-42f0-ab50-2fe8a46dc51a · inbound

Revealing Physical-World Semantic Vulnerabilities: Universal Adversarial Patch for Infrared Vision-Language Models cites this paper.

Revealing Physical-World Semantic Vulnerabilities: Universal Adversarial Patch for Infrared Vision-Language Models Learning Transferable Visual Models From Natural Language Supervision

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-13T19:48:11.661565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T19:43:29.058335Z digest=sha256:e74f31c4a18a502b2731552848d725b0443f5fa7037474377257c8aa789383d3

Observation 2331d670-77a6-47b5-9b75-82dd1a6a3543 · inbound

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders cites this paper.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Learning Transferable Visual Models From Natural Language Supervision

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:561250b8b08090b48cb2959b703f2d3c55f52be8233ae3ec69d1d6d33c3b7ff7

Observation 05dc2e97-096e-4864-a28f-fdcd87c30e82 · inbound

VA-FastNavi-MARL: Real-Time Robot Control with Multimedia-Driven Meta-Reinforcement Learning cites this paper.

VA-FastNavi-MARL: Real-Time Robot Control with Multimedia-Driven Meta-Reinforcement Learning Learning Transferable Visual Models From Natural Language Supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T11:36:21.683988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T11:36:21.683988Z digest=sha256:90a3726a05efca5d715968b399a1d428c249677994cbfed8b1ccc312ac3d7386

Observation 4c4c7c1d-db6e-403f-9f95-13aa5dcf545c · inbound

Evaluation of Embedding-Based and Generative Methods for LLM-Driven Document Classification: Opportunities and Challenges cites this paper.

Evaluation of Embedding-Based and Generative Methods for LLM-Driven Document Classification: Opportunities and Challenges Learning Transferable Visual Models From Natural Language Supervision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T10:23:38.252227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T10:23:38.252227Z digest=sha256:3382450749c8b1613336013c9306ff94fa4aaee8410bcba11f5edec99ea0d2bb

Observation 9dbdcdf6-c129-48d7-b6c7-898fa522092b · inbound

Integration of Object Detection and Small VLMs for Construction Safety Hazard Identification cites this paper.

Integration of Object Detection and Small VLMs for Construction Safety Hazard Identification Learning Transferable Visual Models From Natural Language Supervision

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T19:20:44.287391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-10T19:20:29.951227Z digest=sha256:102b73df5c5948513d87643cf8c0b90e4bb562b5d5d64b70ac5fbd39383a54e2

Observation 3afdcec2-d0d3-4d3f-a307-0b740251f4e9 · inbound

DAT: Dual-Aware Adaptive Transmission for Efficient Multimodal LLM Inference in Edge-Cloud Systems cites this paper.

DAT: Dual-Aware Adaptive Transmission for Efficient Multimodal LLM Inference in Edge-Cloud Systems Learning Transferable Visual Models From Natural Language Supervision

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:10:51.748601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-10T19:18:27.764572Z digest=sha256:eced5622d010d68700493d36d8cc6d22fed8ce304c8f79acde8b881c270b95b1

Observation 2f09bf28-bfdb-4317-b582-fd28b75ff60a · inbound

From Perception to Autonomous Computational Modeling: A Multi-Agent Approach cites this paper.

From Perception to Autonomous Computational Modeling: A Multi-Agent Approach Learning Transferable Visual Models From Natural Language Supervision

Reference 67

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T17:55:41.712110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-10T17:54:40.177049Z digest=sha256:dad11462fd724697d27089db4b0526004eba43bd9c9718ae225f435a99df8756

Observation 4607bcbf-69e7-4931-8059-12283498cde8 · inbound

LLM-Generated Fault Scenarios for Evaluating Perception-Driven Lane Following in Autonomous Edge Systems cites this paper.

LLM-Generated Fault Scenarios for Evaluating Perception-Driven Lane Following in Autonomous Edge Systems Learning Transferable Visual Models From Natural Language Supervision

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:03:25.286603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T22:58:25.053240Z digest=sha256:4fb90cf469d1174e1f87968071bc7771e190bd4c06b0bbb489140ca247ec79e1

Observation 67243991-f637-43e8-98ee-a6992c8f9d3b · inbound

Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models cites this paper.

Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models Learning Transferable Visual Models From Natural Language Supervision

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:51:18.245496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-10T17:25:34.942028Z digest=sha256:5068e48ff9f51e3d39026b660edc2a795deb39cb4ef6f6d9855514254c283699

Observation ec1c05a7-9d05-474d-870d-fd1370205f56 · inbound

ADAPTive Input Training for Many-to-One Pre-Training on Time-Series Classification cites this paper.

ADAPTive Input Training for Many-to-One Pre-Training on Time-Series Classification Learning Transferable Visual Models From Natural Language Supervision

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:25:59.986372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-10T17:11:45.363802Z digest=sha256:18a26e06b6690253225f99eddc64e2dfd622f27df65e33c846ca2339be431dd2

Observation cf40745e-4c45-45e0-ac39-67a57566ab0b · inbound

WildDet3D: Scaling Promptable 3D Detection in the Wild cites this paper.

WildDet3D: Scaling Promptable 3D Detection in the Wild Learning Transferable Visual Models From Natural Language Supervision

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:26:00.377701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-10T17:38:13.336003Z digest=sha256:35ce2a348ba1cf8d6607a1f7825d15c515a325b081648a3f8444a8d5ca20a312

Observation 28539840-fa12-472b-8314-5ff9a8dd276f · inbound

Low-Data Supervised Adaptation Outperforms Prompting for Cloud Segmentation Under Domain Shift cites this paper.

Low-Data Supervised Adaptation Outperforms Prompting for Cloud Segmentation Under Domain Shift Learning Transferable Visual Models From Natural Language Supervision

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T07:41:01.043310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-10T17:02:34.229928Z digest=sha256:24325b739e4848f4dbabaa11aa0117b5b9c6a33b3221a0f84730c63f486f262e

Observation 61734c1e-0ce5-449b-b559-396e1b110e3c · inbound

Few-Shot Contrastive Adaptation for Audio Abuse Detection in Low-Resource Indic Languages cites this paper.

Few-Shot Contrastive Adaptation for Audio Abuse Detection in Low-Resource Indic Languages Learning Transferable Visual Models From Natural Language Supervision

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T06:56:04.171477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-10T17:21:53.195676Z digest=sha256:ddb984993365c2edeb4a6d0f866498ec97d589f947a3c5000c4eabd69fdb5590

Observation 6e82c518-9402-4c02-be65-872dcc0b5386 · inbound

Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection cites this paper.

Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection Learning Transferable Visual Models From Natural Language Supervision

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:35:57.244771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-10T17:07:17.174815Z digest=sha256:cc9974eac97d5007e310f8f8ae2260a8c5a4e93d59f5be6b3e61644e07615863

Observation 4a52da17-99cd-4b01-a3ac-e5e9d121119b · inbound

CWCD: Category-Wise Contrastive Decoding for Structured Medical Report Generation cites this paper.

CWCD: Category-Wise Contrastive Decoding for Structured Medical Report Generation Learning Transferable Visual Models From Natural Language Supervision

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T21:25:48.631995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:28f80b40da07b913b991f090e52ba3efd1bb2e86ad9963fb7cc2676eb09d76f3

Observation 240de0b9-e7de-4a67-9187-5e5b779c4aee · inbound

WARPED: Wrist-Aligned Rendering for Robot Policy Learning from Egocentric Human Demonstrations cites this paper.

WARPED: Wrist-Aligned Rendering for Robot Policy Learning from Egocentric Human Demonstrations Learning Transferable Visual Models From Natural Language Supervision

Reference 140

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T11:01:02.607834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-10T15:14:24.932972Z digest=sha256:ffe812132b8472d45c0ec11fb35cfc2b54aca19e5b3597a712cf78c3bfe7f2f1

Observation 83e66777-da86-4adc-a8ce-093679bde057 · inbound

Improving Layout Representation Learning Across Inconsistently Annotated Datasets via Agentic Harmonization cites this paper.

Improving Layout Representation Learning Across Inconsistently Annotated Datasets via Agentic Harmonization Learning Transferable Visual Models From Natural Language Supervision

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:50:58.266493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-10T16:28:16.315767Z digest=sha256:25f9d681b25cfb0a4f21374138422f768ddd9d2479cccfe87c1c867409eb6b13

Observation eba6f9c0-5eb8-43ea-af35-92ab65aed1a6 · inbound

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models cites this paper.

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models Learning Transferable Visual Models From Natural Language Supervision

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:41:01.935295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-10T15:53:51.162967Z digest=sha256:1a2683cf7011e2c27f4e381827235850ed25ed7617a242f5d86f55acacf8231e

Observation 973cb886-4515-4b1d-be2d-e4c7485e4ff8 · inbound

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models cites this paper.

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models Learning Transferable Visual Models From Natural Language Supervision

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:17:28.774333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T07:16:15.202466Z digest=sha256:80497a38a360ef7ada4c3db6ba9d2112077349cfd0e7e6c11807ec17ed170f40

Observation 4b2834ba-b81f-422c-b9da-58f2caa490c5 · inbound

Bottleneck Tokens for Unified Multimodal Retrieval cites this paper.

Bottleneck Tokens for Unified Multimodal Retrieval Learning Transferable Visual Models From Natural Language Supervision

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T11:01:03.418532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-10T15:14:00.615638Z digest=sha256:d603c7d92e78dced65aad8ebea73de193243e3066a98b9e26b8db915921276e6

Observation b2611979-b265-43e2-b8cb-46beba18dd62 · inbound

Grounded World Model for Semantically Generalizable Planning cites this paper.

Grounded World Model for Semantically Generalizable Planning Learning Transferable Visual Models From Natural Language Supervision

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-10T15:05:32.201428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-10T15:05:29.465402Z digest=sha256:b82d83e11e78715d98ba0a5c97f1efbd17f742060a1b29983039d80e8ceb1f0e

Observation b2cbc247-d5e5-408d-b80e-6e757bfeefad · inbound

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding cites this paper.

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding Learning Transferable Visual Models From Natural Language Supervision

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:31:01.399061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-10T14:50:37.022338Z digest=sha256:d5c63d4e7b3e4ca7f39b82eaacc0904dbeecbf4bde889303bdb16fb1d9a2fb31