Pith. sign in

Paper Citation Record · LEDGER

Cosmos 3: Omnimodal World Models for Physical AI

As of 23 July 2026, this Paper Citation Record lists 15 of 15 outbound references and 27 inbound Pith citation observations for arXiv:2606.02800.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.02800 v4

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T15:08:33.957835Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-20T06:30:07.809122+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-15T10:20:59.147440Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:15:01.137692Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch9

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 90ab2315-cd92-4d4c-b723-4373ad1efe21 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Cosmos 3: Omnimodal World Models for Physical AI PaliGemma: A versatile 3B VLM for transfer

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:46:18.409103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:fff2370bed4462a6fae74026120f492a879576737f87ba4fbf188a437e075d28

Observation 92d4c903-e3f6-41ab-b4b9-d406b0253b4e · outbound

This paper cites Internvla-a1: Unifying understanding, generation and action for robotic manipulation.

Cosmos 3: Omnimodal World Models for Physical AI Internvla-a1: Unifying understanding, generation and action for robotic manipulation

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:18.364550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:06dedee88f4c42ad422f52912e88588b4285151e6fdb6a5e386124fc54561a46

Observation e3df3b5b-4aa8-4deb-9ede-690dc6f95bab · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

Cosmos 3: Omnimodal World Models for Physical AI GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:46:18.382387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:372783da824786a97be1d8a910c9a84694e00dbad3c066be566426fe57b798c1

Observation 7fc89fb4-9072-4458-b829-3c7d59a981c5 · outbound

This paper cites Out of time: Automated lip sync in the wild.

Cosmos 3: Omnimodal World Models for Physical AI Out of time: Automated lip sync in the wild

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T15:08:33.957835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:b40ea2895bae78d2f54a3dc1467de25abe34a75975a91c0f43acea39cb79bb0d

Observation 82577fa1-45ae-4aa8-ba34-435ca39c7be7 · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

Cosmos 3: Omnimodal World Models for Physical AI NVLM: Open Frontier-Class Multimodal LLMs

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:18.387412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:3defb6ed32a72cf2f59935c2f557c73ad7cb26bb2e225432a918fd2244a040dc

Observation b7c293f8-d6ef-4633-9cef-b54d65c05d07 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Cosmos 3: Omnimodal World Models for Physical AI Emerging Properties in Unified Multimodal Pretraining

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:46:18.392456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:8512ced1ac83282ecfb02664a9d9567cf39854402d984cc99183d422fb8331aa

Observation 26399458-3208-4e5f-ad80-20b554940130 · outbound

This paper cites VLMEvalKit: An open-source toolkit for evaluating large multi-modality models.

Cosmos 3: Omnimodal World Models for Physical AI VLMEvalKit: An open-source toolkit for evaluating large multi-modality models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T15:08:33.957835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:9c89fd8a63dd74494e3a19e36a62e0fe19f5014f17bc53c7be6ddd81a5e685d4

Observation a1e8a19a-7040-462b-b9c1-5a55972fb968 · outbound

This paper cites CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models.

Cosmos 3: Omnimodal World Models for Physical AI CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:36:18.020144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:73e61e93945978e27ad0ee91b8a1d895e80cf34771a940198ba246f3a578d5f8

Observation b47562f0-b06b-412d-a181-48b19ee21455 · outbound

This paper cites CameraCtrl: Enabling Camera Control for Text-to-Video Generation.

Cosmos 3: Omnimodal World Models for Physical AI CameraCtrl: Enabling Camera Control for Text-to-Video Generation

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:36:18.034862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:aebf4750746109fec28f9fd5f87dbd79e12fd1de95dc23012b2bbe0ec5a0f524

Observation 3131c469-4361-4749-a281-04a1258d019f · outbound

This paper cites MolmoAct: Action Reasoning Models that can Reason in Space.

Cosmos 3: Omnimodal World Models for Physical AI MolmoAct: Action Reasoning Models that can Reason in Space

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:46:18.358146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:59125e7ddb51d33b51382d6717f4030fe8f4f3b2cd4866d7f49979efc77cc25a

Observation 56a83cb6-2c8d-42b5-86c2-232b9e46d46c · outbound

This paper cites SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes.

Cosmos 3: Omnimodal World Models for Physical AI SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:46:18.396613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:4a2583a708affd2e0689ae2601d2fed1c15061c548939aeb1b13e9fcda5adfa7

Observation 3b55fcdc-a366-48ec-bb80-095cba4609f7 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Cosmos 3: Omnimodal World Models for Physical AI GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:46:18.405483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:c6600d3a1455cea35e86cbe8b44dafe690f3862bd6f3c8ca39be6a3f70907917

Observation 29b28eea-60da-4b26-8ade-f90d4a326b3a · outbound

This paper cites Learning to Act without Actions.

Cosmos 3: Omnimodal World Models for Physical AI Learning to Act without Actions

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:18.377940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:7eacab159268afb4aec7c249d6f22c3168b13f108a42dc72abfb4383cb19760e

Observation 0ed2161f-86dc-48bf-80ec-fcf968de0120 · outbound

This paper cites Video models are zero-shot learners and reasoners.

Cosmos 3: Omnimodal World Models for Physical AI Video models are zero-shot learners and reasoners

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:46:18.400971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:159710ef9f12ffbd0040e486fbdef53848f9aeb060c73b0c64169faf610169fd

Observation 61e6f52a-62a4-41f7-bfa8-7662d99a7d2d · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

Cosmos 3: Omnimodal World Models for Physical AI Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:46:18.371943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:6f2400d89e887f52e75c04e5d153870374c36f4a94bc8897103cba33d58509ba

Pith citing papers

Observation 29cba8cc-5629-4e9f-953f-62913f2fb8f9 · inbound

What Spatial Memory Must Store: Occlusion as the Test for Language-Agent Memory cites this paper.

What Spatial Memory Must Store: Occlusion as the Test for Language-Agent Memory Cosmos 3: Omnimodal World Models for Physical AI

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:37:37.213686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-27T13:46:29.718269Z digest=sha256:b771db42e60832767c53c4927f5fe8f8bc58ee1e5b53bddd2a5bbdea14ccb015

Observation 91adc934-0bf2-40d7-944c-8c6ddbb0fb3a · inbound

ActWorld: From Explorable to Interactive World Model via Action-Aware Memory cites this paper.

ActWorld: From Explorable to Interactive World Model via Action-Aware Memory Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:28:55.397948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-27T01:22:12.771098Z digest=sha256:fbfcac7ceeb28648d24353e01365b7e6e55e53ec9298ebe2eeb17ab4706e92c7

Observation 6ca8f406-2a4b-422f-9777-76ce08f88f85 · inbound

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation cites this paper.

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation Cosmos 3: Omnimodal World Models for Physical AI

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:38:58.754179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-27T00:19:33.170645Z digest=sha256:e9efdcd9c2c3392ec692a680fd8eec43e2a8fe1acdae9723fbe074c99d010ac3

Observation fa8e1f51-2122-4199-8976-eac1433f015c · inbound

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation cites this paper.

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation Cosmos 3: Omnimodal World Models for Physical AI

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:19:13.362817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-26T21:18:07.494794Z digest=sha256:1e9292a203b510d37caaa7ef57ab9205e497e8e0b97671124026b7daf85b5102

Observation b6c19015-8863-419a-a00d-ae261b110023 · inbound

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation cites this paper.

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation Cosmos 3: Omnimodal World Models for Physical AI

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:51.389791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-29T05:06:30.594399Z digest=sha256:8d702df8ce49b3f5707b54eedfc4d1ea0769e83c71f24b52b6c8a7771087a9a7

Observation 4a382069-ca46-4061-8322-4b9cea202198 · inbound

Physics-IQ Verified cites this paper.

Physics-IQ Verified Cosmos 3: Omnimodal World Models for Physical AI

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:29:15.556417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-26T21:13:36.568700Z digest=sha256:67d1a8743c09b240fd72457cac73dd7a4a2f1b9fe497629d0318150568ba6d37

Observation 891ebd45-3d91-4e59-b077-7546724672b3 · inbound

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? cites this paper.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Cosmos 3: Omnimodal World Models for Physical AI

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.503977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:759d389b8ce8e104c63d46c3a114bed7ec2b584113f90ef9ad6fa43397a82da3

Observation dc43e339-80bc-457f-acb7-8a04d4428f56 · inbound

Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation cites this paper.

Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation Cosmos 3: Omnimodal World Models for Physical AI

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:49:41.558859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-26T11:05:44.972128Z digest=sha256:f6334bfc000ef5e5f391089422d38f4d139ceccadc23cf8ed886c18870d2557d

Observation 4fb51621-8fce-4a31-9f5a-9e337858f7a3 · inbound

Critique of Agent Model cites this paper.

Critique of Agent Model Cosmos 3: Omnimodal World Models for Physical AI

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-07-04T11:29:51.076095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-26T07:57:20.830605Z digest=sha256:1bdfc9417bfeed9b2daea391f4009e462542e767cb15b739f4b069e440feecb5

Observation 2d8ece93-2c98-486b-b0a7-44a037d7cca0 · inbound

DiffusionBench: On Holistic Evaluation of Diffusion Transformers cites this paper.

DiffusionBench: On Holistic Evaluation of Diffusion Transformers Cosmos 3: Omnimodal World Models for Physical AI

Reference 212

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T16:59:58.195744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=arxiv_source observed=2026-06-26T00:06:11.951205Z digest=sha256:16fc441aec1f2698c72b12b18f572934d8b002cb7a6b97b2bab425fdb2d8803c

Observation 65459639-2a72-4e6e-b9b7-6ba7fbd8258e · inbound

Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models cites this paper.

Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models Cosmos 3: Omnimodal World Models for Physical AI

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:11.445672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-25T20:57:30.765802Z digest=sha256:48913fd1e20230d234348e657127a6a41e0ab83bc891be2699fbf848391db5ae

Observation e1402191-1b5d-44fd-97dc-807539f2af53 · inbound

Learning Action Priors for Cross-embodiment Robot Manipulation cites this paper.

Learning Action Priors for Cross-embodiment Robot Manipulation Cosmos 3: Omnimodal World Models for Physical AI

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-04T21:00:09.820221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-25T19:09:56.409766Z digest=sha256:deef31db057973e25db7f7ca6c6a93804f89edbc20f32096d05b3130089a047b

Observation 0707065f-fc4d-47f5-9996-73e35bf64de4 · inbound

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation cites this paper.

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:03:56.948658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-29T04:34:38.286863Z digest=sha256:6f7bd428275980a196a7a5eb89bfb22eb96f451c3b379b6a89eb967760bfeae0

Observation 42699cb3-ed34-4566-a74e-45d7db3f8368 · inbound

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis cites this paper.

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis Cosmos 3: Omnimodal World Models for Physical AI

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-30T10:54:35.850307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T10:49:58.019238Z digest=sha256:10c1d8b9b397df2858005efc9251431e142e322d11e8b14bd9242d3d1ba2a636

Observation 8df2960d-1060-4c18-84e1-97ccce87eeca · inbound

Mural: Transferring LLM knowledge to image generation via Mixture-of-Transformers cites this paper.

Mural: Transferring LLM knowledge to image generation via Mixture-of-Transformers Cosmos 3: Omnimodal World Models for Physical AI

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T09:24:32.131834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T09:23:59.752793Z digest=sha256:894060c24f6e4e0860b855f1d5f18dff63082db396f4b0a3566eeb8561907d02

Observation 8df1cee3-7063-430e-9ec5-4fda6ccc8145 · inbound

World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration cites this paper.

World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration Cosmos 3: Omnimodal World Models for Physical AI

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:25:41.706604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T05:31:43.678762Z digest=sha256:fc5b3138f378569fe68b5f67db4f334e7e17f0ec39a8b24fca9f0c3b903d6b8b

Observation 5aae7d9a-c867-495b-95f1-9df375d77f2f · inbound

World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration cites this paper.

World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration Cosmos 3: Omnimodal World Models for Physical AI

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-15T10:20:59.147440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:20:59.147440Z digest=sha256:083e10bc76949e95194d099c48798291a9a35dbad8b784128a2fc11ed80c6ad2

Observation 595f0b64-ad7b-46dc-a55b-0de06bfff587 · inbound

ROSA: A Robotics Foundation Model Serving System for Robot Factories cites this paper.

ROSA: A Robotics Foundation Model Serving System for Robot Factories Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:16:52.666121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-02T11:11:39.584755Z digest=sha256:81a1df7b4c05dfc1b16b40f8b2347a956166b7138c7e59c227a725ca2d9912dd

Observation 6e6bf47c-766d-4292-bb0c-2aeecf7e3122 · inbound

DynaWM: A Base-VLA-Guided World Foundation Model for Moving-Object Manipulation cites this paper.

DynaWM: A Base-VLA-Guided World Foundation Model for Moving-Object Manipulation Cosmos 3: Omnimodal World Models for Physical AI

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:40.344868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:40.344868Z digest=sha256:06bd32ef291cbebbb9390fcb753e74f991aea70a68760b97d146835564dd6718

Observation e815add8-2379-4d1f-97f4-1b2bbe946ad3 · inbound

GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation cites this paper.

GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T08:04:48.963890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:04:48.963890Z digest=sha256:fb98fbd075d7613e18f44bb6626ab535a97122c6b814a2a845396aa987ee3180

Observation 2e3ebbcd-57e3-4dbe-8636-e49b96634871 · inbound

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI cites this paper.

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI Cosmos 3: Omnimodal World Models for Physical AI

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:57.396710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:57.396710Z digest=sha256:efce88316aaed4d18cbddab622db5d28de1fc0ecfc554a7612387b5148b56478

Observation 340b9926-bd9a-4491-b57a-94fc9fb807a6 · inbound

A Definition and Roadmap for World Models cites this paper.

A Definition and Roadmap for World Models Cosmos 3: Omnimodal World Models for Physical AI

Reference 128

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:44.336323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=arxiv_source observed=2026-07-08T07:10:33.826140Z digest=sha256:0ea73da498e94e52a747f7b8763a0ed561f3d01e6ff0c65e6561a5abc358a7d3

Observation 5cfc9951-4d25-4d42-a0ba-7eb2e0803ad2 · inbound

From Foundation to Application: Improving VLA Models in Practice cites this paper.

From Foundation to Application: Improving VLA Models in Practice Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.485926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:031bab9e46cb253a0f34fa5d48726710e2cefb28b391a9f75706a45be3a677bf

Observation fea50311-ff49-4e5f-9c6d-776b88c0059e · inbound

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time cites this paper.

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time Cosmos 3: Omnimodal World Models for Physical AI

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T06:48:14.554799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:48:14.554799Z digest=sha256:e3002f4c324a127fd53973d7c3286c6d75c067725db883906ec9c2e13db2c215

Observation d70e5411-8c99-4851-8678-3fdf7bf4fa3f · inbound

Infinite Worlds with Versatile Interactions cites this paper.

Infinite Worlds with Versatile Interactions Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-09T07:56:04.717628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-09T07:51:36.802801Z digest=sha256:6323ad172e44eaba89ea5a2096a52fd2d2244ee0238494103216aa0838d31289

Observation ad2ace2e-2a21-4fd2-b5d8-120e307a1492 · inbound

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence cites this paper.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Cosmos 3: Omnimodal World Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.157813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:04d20a5d0cd54657ca20ac9ed60c41376a3b0e089febada67cda6a6ba5bcae49

Observation 25d1cb50-f7ab-4983-8cba-96e8942e7e39 · inbound

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model cites this paper.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:85466e3bdecc3deb91c0621e1465ecee05e8a9155794d148c7154c1919502d08