REVIEW 3 major objections 5 minor 162 references
Robot policy evaluators succeed by staying action-faithful over long horizons, not by looking more photorealistic.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 08:03 UTC pith:KFY7ODQ4
load-bearing objection Solid large-scale empirical roadmap for world models as robot policy evaluators; the 14.9% headline is on a chosen diagnostic average, not quantified closed-loop ranking ρ. the 3 major comments →
GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
A world model is a reliable robot-policy evaluator only when its closed-loop rollouts stay action-faithful over long horizons and therefore reproduce the same success/failure ranking that real robots produce; short-term photorealism is secondary, and the decisive levers are balanced physical-plus-robot data, spatially aligned action control, persistent multi-scale memory, and post-training aimed at evaluator agreement rather than generic video quality.
What carries the argument
WMBench: a paired real/world-model rollout benchmark whose primary target is ranking correlation between real-world and world-model success rates (and the related ordinal WMES score), used to isolate which metrics, data mixes, and architectures actually predict real policy outcomes.
Load-bearing premise
Agreement on WMBench’s eight held-out manipulation families, still drawn from the same platforms, cameras, and task distribution, is enough to claim that a world model will rank novel policies reliably under broader initial states, embodiments, and contact-rich failures.
What would settle it
Take a set of policies whose real-robot success ranking is known on tasks outside WMBench’s eight families or on a different embodiment; if the world model’s closed-loop ranking correlation collapses while short-horizon visual metrics remain high, the central claim that long-horizon action faithfulness on this benchmark is sufficient fails.
If this is right
- Policy developers can replace a large fraction of hardware rollouts with closed-loop world-model evaluation once the model is scored by ranking agreement rather than frame beauty.
- Metric suites that reward static or action-ignorant videos will systematically promote weak evaluators and should be dropped from evaluator leaderboards.
- Training recipes must mix broad physical video with robot data; robot-only fine-tuning improves embodiment look but can erase the priors needed for reliable evaluation.
- Spatially aligned control maps plus hierarchical memory become standard requirements for any video world model intended as a policy surrogate.
- Open release of the benchmark, models, and annotation toolkit lets the community iterate on evaluator design the way language-model groups iterate on digital suites.
Where Pith is reading between the lines
- The same long-horizon action-faithfulness test could become a filter for world models used as data engines or planners, not only as offline evaluators.
- If VLM outcome labeling stays within a few percent of human WMES at method level, most future evaluator leaderboards can run without exhaustive human annotation.
- Contact-rich failure modes still show optimistic bias; hybrid world models that add explicit contact or force state may be the next necessary step beyond pure video.
- Once ranking correlation is the accepted target, sim-to-real gaps in classical simulators can be measured against the same yardstick rather than against visual fidelity alone.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies world models as surrogate evaluators of robot policies, arguing that real-robot evaluation is the bottleneck for embodied foundation models. It introduces WMBench (paired teleoperation and policy-rollout data over eight manipulation task families), analyzes seven video world models, four action encodings, and 324k+ annotated rollouts (including CVPR 2026 challenge submissions), and reports three design insights: long-horizon action-faithful consistency dominates short-term visual realism; pretraining gains require balancing general physical priors with robot controllability; and action interface, memory, and evaluator-oriented post-training strongly affect real-world alignment. These are instantiated in GigaWorld-1 (Wan-based, ~13k hours multi-source data, pixel-aligned EE/ray control, hierarchical memory, progressive training), which improves a six-metric average by 14.9% over Wan 2.2 5B under matched post-training (Table 9). Code, models, and data are released.
Significance. If the design claims hold, the work is a substantial contribution to scalable robot policy evaluation: it reframes world models as external policy evaluators rather than only data engines or planners, provides a large paired real/sim benchmark and metric analysis (Figs. 4–5, WMES), and ships a concrete open roadmap (data mixture, spatially aligned control, memory, distillation) with reproducible artifacts. The controlled ablations (Tables 2–4, 9) and community-scale annotation are genuine strengths relative to prior proof-of-concept evaluator papers. The practical value depends on whether closed-loop ranking fidelity—not only diagnostic video metrics—is demonstrated at the same scale as the headline gains.
major comments (3)
- [Sec. 3 Eq. (4); Sec. 6.5.1 Table 9; Sec. 6.5.5 Figs. 16–17] Sec. 3, Eq. (4) defines the primary evaluator target as ranking/success agreement ρ = Corr(S_real(π), S_wm(π)) across policies. The abstract, intro, and Sec. 6.5.1 instead headline a 14.9% gain of GigaWorld-1-Plus over Wan 2.2 5B on a paper-chosen six-metric average (Aesthetic, Image, JEPA, Semantic, Subject, Trajectory; Table 9: 0.6834 vs 0.5948). Those diagnostics are justified by correlation with human WMES on challenge rollouts (Findings 1–2, Figs. 4–5), but they are not ρ. Closed-loop evidence in Sec. 6.5.5 is limited to task-level success-rate scatter and Gen−Real bias bars on four tasks/subtasks (Figs. 16–17), without a quantified ρ (or Kendall/Spearman ranking) over a multi-checkpoint policy set for GigaWorld-1 vs baselines. Please report ρ (with CIs) under the same closed-loop protocol for the main models, or reframe the 14.9% claim so it is not presented as the primary evaluato
- [Sec. 6.5.5; Fig. 17; Finding 3 / Table 4] Sec. 6.5.5 and Fig. 17 acknowledge residual optimistic bias on contact-sensitive failures (e.g., pour/press subtasks). Because the paper’s own Finding 3 and Table 4 argue that long-horizon action-faithful consistency—not short-horizon visual quality—dominates evaluator reliability, the closed-loop calibration gap is load-bearing. Either quantify how often GigaWorld-1 flips real success/failure relative to baselines (confusion matrices or per-subtask agreement rates on the full WMBench closed-loop set), or temper claims that the model is “specially optimized for policy evaluation” until failure-mode calibration is measured at the same scale as the diagnostic average.
- [Sec. 4.3; Findings 4–5; Table 1; Table 9] WMES and several diagnostic metrics (Perspectivity, Instruction Following, Interaction Quality, Semantic Alignment) rely on VLM judges (Sec. 4.3; Finding 4–5; Table 1). The LoRA VLM is trained to predict human WMES with score-token weight 8.0, then used for scalable outcome assessment. Metric selection for Table 9 is guided by correlation with that same WMES family. This is not fatal circularity—human WMES and real success labels remain external—but it creates mild self-reinforcement risk if models are tuned toward VLM-preferred appearance/geometry. Please report (i) human-only vs VLM-only ranking of the main models on a held-out subset, and (ii) sensitivity of the six-metric average and any ρ to excluding VLM-judged diagnostics.
minor comments (5)
- [Fig. 1; Abstract; Sec. 4.1; Table 6] Fig. 1 and abstract claim “324,000+ analyzed rollouts” and “12K+ hours training data”; Table 6 totals ~12,980 hours. Align the rounded figures and clarify whether 324k counts segments or full closed-loop episodes (Sec. 4.1 says segments chained into episodes of 20–30 segments).
- [Table 3; Finding 9] Table 3 ranks control interfaces on Trajectory Accuracy and motion metrics for Wan 2.1 1.3B only. A short note on whether channel-concat remains best for the 5B backbone (GigaWorld-1-Plus) would strengthen Finding 9.
- [Sec. 4.1; Sec. 6.5.4; Sec. 7] Sec. 4.1 train/test split is episode-disjoint within the same eight task families and platforms. Sec. 7 correctly flags limited coverage of mobile/dexterous/safety-critical settings; a brief explicit statement in Sec. 6.5 that OOD claims (Fig. 15) are appearance/content shifts, not embodiment or policy-family shifts, would prevent over-reading.
- [Fig. 2; Fig. 3] Typo in Fig. 3 caption/step labels: “Train Wodel Model” should be “Train World Model.” Several figure panels (e.g., Fig. 2 Chinese annotation fragment in the source) should be cleaned for the camera-ready version.
- [Finding 6 footnote; Sec. 7] Cosmos-3 is mentioned as planned but unavailable (footnote in Finding 6). Either update the comparison if multiview access is obtained, or move the note to a single limitations paragraph to avoid dangling promises.
Circularity Check
No load-bearing circular derivation: claims rest on external real-robot pairings and held-out metrics; only mild metric-suite selection via human WMES.
specific steps
-
other
[Sec. 5.1 Findings 1–2; Sec. 6.5.1 / Table 9; Eq. 4 vs. reported AVG]
"evaluator quality is dominated by long-horizon, action-faithful rollout consistency rather than short-term visual realism... GigaWorld-1-Plus improves the average score by ... 14.9% over Wan 2.2 5B... we retain six core evaluator-relevant metric(Aesthetic Quality, Image Quality, JEPA Similarity, Semantic Alignment, Subject Consistency, and Trajectory Accuracy)... Visual Fidelity has the highest correlation (ρ=0.78)... Subject Consistency (ρ=0.88) and Perspectivity (ρ=0.86) are the strongest predictors"
The paper’s stated primary target is ρ=Corr(S_real,S_wm) (Eq. 4), but the headline 14.9% is the mean of six automatic metrics pre-selected because they correlate with human WMES on the same challenge ecosystem. That makes the reported “evaluator-alignment” gain partly an improvement on a human-rubric-derived metric suite rather than a direct closed-loop ranking result. Mild self-reinforcement of the evaluation criterion, not a by-construction identity between fit and prediction.
full rationale
This is an empirical systems paper, not a first-principles derivation. The primary evaluator target (Eq. 4) is ranking/success agreement between world-model and real-robot outcomes—an external quantity. WMBench uses episode-disjoint held-out trajectories with real teleoperation and policy rollouts; the 324k annotated segments and human WMES are independent labels, not fitted parameters renamed as predictions. Ablations (control interfaces, memory, data mixtures) and the Table 9 comparison (matched post-training of multiple backbones) measure independently computed diagnostics (JEPA, NDTW trajectory accuracy, etc.). Self-citations (GigaBrain, GigaWorld-0, etc.) supply data sources and related work, not uniqueness theorems that force the result. The only mild circularity risk is operational: Findings 1–2 select the six-metric AVG reported as “evaluator-alignment” because those metrics correlate with human WMES on challenge rollouts, and a LoRA VLM is trained to predict that same WMES for scalable closed-loop scoring—so the reported 14.9% is alignment with a paper-chosen human-rubric proxy, not a direct measurement of ρ. That is metric-selection self-reinforcement, not definitional circularity or a fitted input called a prediction; the model scores are not forced by construction. Score 1 reflects that minor operational loop without elevating it to a load-bearing circular chain.
Axiom & Free-Parameter Ledger
free parameters (6)
- Six-metric evaluator average weights (equal mean of Aesthetic, Image, JEPA, Semantic, Subject, Trajectory)
- VLM score-token loss weight 8.0 (format 1.0, rationale min 0.05)
- Data mixture composition (PhysData / AgiBot / GigaData hours and filters)
- Video quality / motion filter thresholds (τ_img, τ_aes, τ_jump, τ_static, τ_motion, τ_jerk)
- LoRA ranks/alphas and stage learning rates (e.g., Stage1 LR 5e-5, Stage2 1e-4, ranks 128/256)
- Hierarchical memory partition (short/mid/long + first-frame anchor) and generation window length
axioms (5)
- domain assumption Agreement between world-model and real-world policy success/ranking is the primary definition of evaluator quality.
- domain assumption Video diffusion backbones with action conditioning can capture decision-relevant physical dynamics for closed-loop manipulation evaluation.
- domain assumption Human WMES ordinal labels (0–3) and LoRA-tuned VLM scores are valid ground truth for ranking world models as evaluators.
- ad hoc to paper Episode-disjoint train/test splits within the same eight task families suffice to measure generalization for evaluator claims.
- standard math Standard flow-matching / diffusion training and autoregressive windowing preserve enough dynamics for long-horizon evaluation when memory and control are added.
invented entities (3)
-
WMBench
independent evidence
-
WMES (World Model as Evaluator Score)
no independent evidence
-
GigaWorld-1 (Nano/Plus)
independent evidence
read the original abstract
Evaluating embodied robot foundation models remains a critical bottleneck; unlike large language models efficiently assessed via digital benchmarks, robotic policies require slow, costly real-world rollouts limited by hardware and human supervision, which has driven interest in world models as surrogate policy evaluators, yet the key properties that make a world model reliable for policy assessment remain poorly understood. This work presents a systematic study of world models for robotic policy evaluation and introduces WMBench, a benchmark constructed from real-robot teleoperation data and matched policy rollouts covering diverse manipulation tasks to enable controlled comparisons across model families, action encodings, rollout horizons, and evaluation metrics. Using WMBench, we analyze 7 video world models, 4 action representation schemes, and over 324,000 simulated policy rollouts paired with real robot executions, further enriching our analysis with large-scale community submissions from the CVPR 2026 GigaBrain Challenge, curated synthetic trajectories, and a training videos spanning more than 12,000 hours. Our experiments deliver three core insights: evaluator quality is dominated by long-horizon, action-faithful rollout consistency rather than short-term visual realism; pretraining gains stem not only from data scale but from balancing general world knowledge with robot-specific controllability; and architectural choices including action encoding, memory design, and evaluator-focused post-training strongly determine alignment with real-world robot behavior. Drawing on these results, we derive a practical design roadmap and realize it in \textit{GigaWorld-1}, a world model specially optimized for policy evaluation, and we fully release our code, models, datasets, and toolkits to advance scalable evaluation research for embodied foundation models.
Figures
Reference graph
Works this paper leans on
-
[1]
Cosmos 3: Omnimodal world models for physical ai.arXiv preprint arXiv:2606.02800, 2026
Niket Agarwal, Arslan Ali, Jon Allen, Martin Antolini, Adeline Aubame, Alisson Azzolini, Junjie Bai, Maciej Bala, Yogesh Balaji, Josh Bapst, et al. Cosmos 3: Omnimodal world models for physical ai.arXiv preprint arXiv:2606.02800, 2026. 2, 3, 4, 11, 30
Pith/arXiv arXiv 2026
-
[2]
Hassan Abu Alhaija, Jose Alvarez, Maciej Bala, Tiffany Cai, Tianshi Cao, Liz Cha, Joshua Chen, Mike Chen, Francesco Ferroni, Sanja Fidler, et al. Cosmos-transfer1: Conditional world generation with adaptive multimodal control.arXiv preprint arXiv:2503.14492, 2025. 3
Pith/arXiv arXiv 2025
-
[3]
World simulation with video foundation models for physical ai.arXiv preprint arXiv:2511.00062, 2025
Arslan Ali, Junjie Bai, Maciej Bala, Yogesh Balaji, Aaron Blakeman, Tiffany Cai, Jiaxin Cao, Tianshi Cao, Elizabeth Cha, Yu-Wei Chao, et al. World simulation with video foundation models for physical ai.arXiv preprint arXiv:2511.00062, 2025. 11
Pith/arXiv arXiv 2025
-
[4]
Pranav Atreya, Karl Pertsch, Tony Lee, Moo Jin Kim, Arhan Jain, Artur Kuramshin, Clemens Eppner, Cyrus Neary, Edward Hu, Fabio Ramos, et al. Roboarena: Distributed real-world evaluation of generalist robot policies.arXiv preprint arXiv:2506.18123, 2025. 3
arXiv 2025
-
[5]
Hyperreel: High-fidelity 6-dof video with ray-conditioned sampling
Benjamin Attal, Jia-Bin Huang, Christian Richardt, Michael Zollhoefer, Johannes Kopf, Matthew O’Toole, and Changil Kim. Hyperreel: High-fidelity 6-dof video with ray-conditioned sampling. InCVPR, pages 16610–16620, 2023. 19
2023
-
[7]
Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025
Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junyang Lin, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan ...
Pith/arXiv arXiv 2025
-
[9]
Qwen2.5-vl technical report.arXiv preprint arXiv:2502.13923,
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-vl technical report.ar...
-
[10]
7 32 GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation
-
[11]
Navigation world models
Amir Bar, Gaoyue Zhou, Danny Tran, Trevor Darrell, and Yann LeCun. Navigation world models. In CVPR, pages 15791–15801, 2025. 2
2025
-
[12]
V-JEPA: Latent video prediction for visual representation learning, 2024
Adrien Bardes, Quentin Garrido, Jean Ponce, Xinlei Chen, Michael Rabbat, Yann LeCun, Mido Assran, and Nicolas Ballas. V-JEPA: Latent video prediction for visual representation learning, 2024. URL https://openreview.net/forum?id=WFYbBOEOtv. 7
2024
-
[13]
Motus: A unified latent action world model.arXiv preprint arXiv:2512.13030, 2025
Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, Shuhe Huang, Haitian Liu, Ruowen Zhao, Yao Feng, Chendong Xiang, Yinze Rong, et al. Motus: A unified latent action world model.arXiv preprint arXiv:2512.13030, 2025. 2, 3
Pith/arXiv arXiv 2025
-
[14]
Johan Bjorck, Fernando Castañeda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, et al. Gr00t n1: An open foundation model for generalist humanoid robots.arXiv preprint arXiv:2503.14734, 2025
Pith/arXiv arXiv 2025
-
[15]
Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, et al. pi0: A vision-language-action flow model for general robot control.arXiv preprint arXiv:2410.24164, 2024. 2
Pith/arXiv arXiv 2024
-
[16]
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets.arXiv preprint arXiv:2311.15127, 2023. 2, 11
Pith/arXiv arXiv 2023
-
[17]
Tiny autoencoder for stable diffusion.https://github.com/madebyollin/taesd,
Ollin Boer Bohan. Tiny autoencoder for stable diffusion.https://github.com/madebyollin/taesd,
-
[18]
GitHub repository. 24
-
[19]
Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems
Qingwen Bu, Jisong Cai, Li Chen, Xiuqi Cui, Yan Ding, Siyuan Feng, Xindong He, Xu Huang, et al. Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems. In IROS, 2025. 15
2025
-
[20]
Sam 3: Segment anything with concepts, 2025
Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Didac Suris, Chai- tanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, Jie Lei, Tengyu Ma, Baishan Guo, Arpit Kalla, Markus Marks, Joseph Greer, Meng Wang, Peize Sun, Roman Rädle, Triantafyllos Afouras, Effrosyni Mavroudi, Katherine Xu, Tsung-Han Wu, Yu Zhou, Lilian...
Pith/arXiv arXiv 2025
-
[21]
Emerging properties in self-supervised vision transformers, 2021
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers, 2021. URLhttps://arxiv.org/ abs/2104.14294. 7
Pith/arXiv arXiv 2021
-
[22]
Worldvla: Towards autoregressive action world model.arXiv preprint arXiv:2506.21539, 2025
Jun Cen, Chaohui Yu, Hangjie Yuan, Yuming Jiang, Siteng Huang, Jiayan Guo, Xin Li, Yibing Song, Hao Luo, Fan Wang, et al. Worldvla: Towards autoregressive action world model.arXiv preprint arXiv:2506.21539, 2025. 2
Pith/arXiv arXiv 2025
-
[23]
Reflect-r1: Evidence-driven reflection for self-correction in long video understanding, 2026
Shuimu Chen, Yuteng Chen, Yuanshen Guan, Zebang Cheng, Zeyu Zhang, Shengqian Qin, Bin Xia, Jiaran Li, Wenming Yang, and Fei Ma. Reflect-r1: Evidence-driven reflection for self-correction in long video understanding, 2026. URLhttps://arxiv.org/abs/2606.27922. 17
Pith/arXiv arXiv 2026
-
[24]
Tianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai, Yibin Liu, Qiwei Liang, Zixuan Li, Xianliang Lin, Yiheng Ge, Zhenyu Gu, et al. Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation.arXiv preprint arXiv:2506.18088,
-
[25]
2, 4 33 GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation
-
[26]
Yuzhi Chen, Ronghan Chen, Dongjie Huo, Yandan Yang, Dekang Qi, Haoyun Liu, Tong Lin, Shuang Zeng, Junjin Xiao, Xinyuan Chang, et al. Abot-physworld: Interactive world foundation model for robotic manipulation with physics alignment.arXiv preprint arXiv:2603.23376, 2026. 2
arXiv 2026
-
[27]
Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021. 2
Pith/arXiv arXiv 2021
-
[28]
Introducing veo 3, our video generation model with expanded creative controls – including native audio and extended videos
Google DeepMind. Introducing veo 3, our video generation model with expanded creative controls – including native audio and extended videos. [Online], 2025. URLhttps://deepmind.google/models/ veo/. 3
2025
-
[29]
Zhehao Dong, Xiaofeng Wang, Zheng Zhu, Yirui Wang, Yang Wang, Yukun Zhou, Boyuan Wang, Chaojun Ni, RunqiOuyang, WenkangQin, etal. Emma: Generalizingreal-worldrobotmanipulationviagenerative visual transfer.arXiv preprint arXiv:2509.22407, 2025. 3
arXiv 2025
-
[30]
Liaoyuan Fan, Zetian Xu, Chen Cao, Wenyao Zhang, Mingqi Yuan, and Jiayu Chen. Aim: Intent-aware unified world action modeling with spatial value maps.arXiv preprint arXiv:2604.11135, 2026. 3
Pith/arXiv arXiv 2026
-
[31]
Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. InCVPR, pages 24108–24118, 2025. 2
2025
-
[32]
Xiao Fu, Xintao Wang, Xian Liu, Jianhong Bai, Runsen Xu, Pengfei Wan, Di Zhang, and Dahua Lin. Learning video generation for robotic manipulation with collaborative trajectory control.arXiv preprint arXiv:2506.01943, 2025. 19
arXiv 2025
-
[33]
Yu Gao, Haoyuan Guo, Tuyen Hoang, Weilin Huang, Lu Jiang, Fangyuan Kong, Huixia Li, Jiashi Li, Liang Li, Xiaojie Li, et al. Seedance 1.0: Exploring the boundaries of video generation models.arXiv preprint arXiv:2506.09113, 2025. 3
Pith/arXiv arXiv 2025
-
[34]
Genesis: Multimodal driving scene generation with spatio-temporal and cross-modal consistency.NeurIPS, 38:60431–60455, 2026
Xiangyu Guo, Zhanqian Wu, Kaixin Xiong, Ziyang Xu, Lijun Zhou, Gangwei Xu, Shaoqing Xu, Haiyang Sun, Bing Wang, Guang Chen, et al. Genesis: Multimodal driving scene generation with spatio-temporal and cross-modal consistency.NeurIPS, 38:60431–60455, 2026. 3
2026
-
[35]
Yanjiang Guo, Lucy Xiaoyang Shi, Jianyu Chen, and Chelsea Finn. Ctrl-world: A controllable generative world model for robot manipulation.arXiv preprint arXiv:2510.10125, 2025. 2, 3, 4
Pith/arXiv arXiv 2025
-
[36]
Ltx-video: Realtime video latent diffusion.arXiv preprint arXiv:2501.00103, 2024
Yoav HaCohen, Nisan Chiprut, Benny Brazowski, Daniel Shalem, Dudu Moshe, Eitan Richardson, Eran Levin, Guy Shiran, Nir Zabari, Ori Gordon, et al. Ltx-video: Realtime video latent diffusion.arXiv preprint arXiv:2501.00103, 2024. 2, 11
Pith/arXiv arXiv 2024
-
[37]
Pre-trained video generative models as world simulators
Haoran He, Yang Zhang, Liang Lin, Zhongwen Xu, and Ling Pan. Pre-trained video generative models as world simulators. InAAAI, volume 40, pages 4645–4653, 2026. 2
2026
-
[38]
Xianglong He, Chunli Peng, Zexiang Liu, Boyang Wang, Yifan Zhang, Qi Cui, Fei Kang, Biao Jiang, Mengyin An, Yangyang Ren, et al. Matrix-game 2.0: An open-source real-time and streaming interactive world model.arXiv preprint arXiv:2508.13009, 2025. 3
Pith/arXiv arXiv 2025
-
[39]
Ziheng He, Yixiang Chen, Ning Yang, Zhanqian Wu, Qisen Ma, Yuan Xu, Jiabing Yang, Peiyan Li, Xiangnan Wu, Xiaofeng Wang, et al. Skip: Sparse keyframe interpolation paradigm for efficient embodied world models.arXiv preprint arXiv:2606.00664, 2026. 3
Pith/arXiv arXiv 2026
-
[40]
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. InICLR, 2020. 2 34 GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation
2020
-
[41]
Gans trained by a two time-scale update rule converge to a local nash equilibrium.NeurIPS, 2017
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium.NeurIPS, 2017. 8
2017
-
[42]
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. InICLR, 2023. 3
2023
-
[43]
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg. Ruler: What’s the real context size of your long-context language models?arXiv preprint arXiv:2404.06654, 2024. 2
Pith/arXiv arXiv 2024
-
[44]
Lora: Low-rank adaptation of large language models
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. Lora: Low-rank adaptation of large language models. InICLR, 2021. 18
2021
-
[45]
Self forcing: Bridging the train-test gap in autoregressive video diffusion.NeurIPS, 38:167283–167308, 2026
Xun Huang, Zhengqi Li, Guande He, Mingyuan Zhou, and Eli Shechtman. Self forcing: Bridging the train-test gap in autoregressive video diffusion.NeurIPS, 38:167283–167308, 2026. 3
2026
-
[46]
Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, et al.𝜋0.5: a vision-language-action model with open-world generalization.arXiv preprint arXiv:2504.16054, 2025. 2
Pith/arXiv arXiv 2025
-
[47]
Rlbench: The robot learning benchmark & learning environment.IEEE Robotics and Automation Letters, 5(2):3019–3026, 2020
Stephen James, Zicong Ma, David Rovick Arrojo, and Andrew J Davison. Rlbench: The robot learning benchmark & learning environment.IEEE Robotics and Automation Letters, 5(2):3019–3026, 2020. 2
2020
-
[48]
Galaxea open-world dataset and g0 dual-system vla model.arXiv preprint arXiv:2509.00576, 2025
Tao Jiang, Tianyuan Yuan, Yicheng Liu, Chenhao Lu, Jianning Cui, Xiao Liu, Shuiqi Cheng, Jiyang Gao, Huazhe Xu, and Hang Zhao. Galaxea open-world dataset and g0 dual-system vla model.arXiv preprint arXiv:2509.00576, 2025. 2, 15
Pith/arXiv arXiv 2025
-
[49]
Yuxin Jiang, Shengcong Chen, Siyuan Huang, Liliang Chen, Pengfei Zhou, Yue Liao, Xindong He, Chiming Liu, Hongsheng Li, Maoqing Yao, et al. Enerverse-ac: Envisioning embodied environments with action condition.arXiv preprint arXiv:2505.09723, 2025. 2
Pith/arXiv arXiv 2025
-
[50]
Swe-bench: Can language models resolve real-world github issues? InICLR, volume 2024, pages 54107–54157, 2024
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. Swe-bench: Can language models resolve real-world github issues? InICLR, volume 2024, pages 54107–54157, 2024. 2
2024
-
[51]
Bingyi Kang, Yang Yue, Rui Lu, Zhijie Lin, Yang Zhao, Kaixin Wang, Gao Huang, and Jiashi Feng. How far is video generation from world model: A physical law perspective.arXiv preprint arXiv:2411.02385,
-
[52]
Musiq: Multi-scale image quality transformer
Junjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar, and Feng Yang. Musiq: Multi-scale image quality transformer. InICCV, pages 5148–5157, October 2021. 7
2021
-
[53]
Openvla: Anopen-sourcevision-language-action model.arXiv preprint arXiv:2406.09246, 2024
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov,EthanFoster,GraceLam,PannagSanketi,etal. Openvla: Anopen-sourcevision-language-action model.arXiv preprint arXiv:2406.09246, 2024. 2
Pith/arXiv arXiv 2024
-
[54]
Moo Jin Kim, Yihuai Gao, Tsung-Yi Lin, Yen-Chen Lin, Yunhao Ge, Grace Lam, Percy Liang, Shuran Song, Ming-Yu Liu, Chelsea Finn, et al. Cosmos policy: Fine-tuning video models for visuomotor control and planning.arXiv preprint arXiv:2601.16163, 2026. 2, 3
Pith/arXiv arXiv 2026
-
[55]
Aesthetic predictor, 2022
LAION-AI. Aesthetic predictor, 2022. URLhttps://github.com/LAION-AI/aesthetic-predictor. Accessed: 2024. 7
2022
-
[56]
Xiaolei Lang, Yang Wang, Yukun Zhou, Chaojun Ni, Kerui Li, Jiagang Zhu, Tianze Liu, Jiajun Lv, Xingxing Zuo, Yun Ye, et al. Vag: Dual-stream video-action generation for embodied data synthesis.arXiv preprint arXiv:2604.09330, 2026. 2 35 GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation
Pith/arXiv arXiv 2026
-
[57]
HuashuoLei,WenxuanSong,HuaruiZhang,JieyuanPei,JiayiChen,HaodongYan,HanZhao,Pengxiang Ding, Zhipeng Zhang, Lida Huang, et al. Robomemarena: A comprehensive and challenging robotic memory benchmark.arXiv preprint arXiv:2605.10921, 2026. 4
Pith/arXiv arXiv 2026
-
[58]
HaoyunLi, IvanZhang, RunqiOuyang, XiaofengWang, ZhengZhu, ZhiqinYang, ZhentaoZhang, Boyuan Wang, Chaojun Ni, Wenkang Qin, et al. Mimicdreamer: Aligning human and robot demonstrations for scalable vla training.arXiv preprint arXiv:2509.22199, 2025. 3
arXiv 2025
-
[59]
Kerui Li, Zhe Jing, Xiaofeng Wang, Zheng Zhu, Yukun Zhou, Guan Huang, Dongze Li, Qingkai Yang, and Huaibo Huang. Stableidm: Stabilizing inverse dynamics model against manipulator truncation via spatio-temporal refinement.arXiv preprint arXiv:2604.17887, 2026. 3
Pith/arXiv arXiv 2026
-
[60]
Evaluatingreal-worldrobotmanipulationpoliciesinsimulation
Xuanlin Li, Kyle Hsu, Jiayuan Gu, Karl Pertsch, Oier Mees, Homer Rich Walke, Chuyuan Fu, Ishikaa Lunawat,IsabelSieh,SeanKirmani,etal. Evaluatingreal-worldrobotmanipulationpoliciesinsimulation. arXiv preprint arXiv:2405.05941, 2024. 2
Pith/arXiv arXiv 2024
-
[61]
Worldeval: World model as real-world robot policies evaluator.arXiv preprint arXiv:2505.19017, 2025
Yaxuan Li, Yichen Zhu, Junjie Wen, Chaomin Shen, and Yi Xu. Worldeval: World model as real-world robot policies evaluator.arXiv preprint arXiv:2505.19017, 2025. 2
Pith/arXiv arXiv 2025
-
[62]
Yue Liao, Pengfei Zhou, Siyuan Huang, Donglin Yang, Shengcong Chen, Yuxin Jiang, Yue Hu, Jingbin Cai, Si Liu, Jianlan Luo, et al. Genie envisioner: A unified world foundation platform for robotic manipulation.arXiv preprint arXiv:2508.05635, 2025. 19
Pith/arXiv arXiv 2025
-
[63]
Depth anything 3: Recovering the visual space from any views.arXiv preprint arXiv:2511.10647,
Haotong Lin, Sili Chen, Junhao Liew, Donny Y Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang. Depth anything 3: Recovering the visual space from any views.arXiv preprint arXiv:2511.10647,
-
[64]
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. Truthfulqa: Measuring how models mimic human falsehoods. InProceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers), pages 3214–3252, 2022. 2
2022
-
[65]
Libero: Bench- marking knowledge transfer for lifelong robot learning.NeurIPS, 36:44776–44791, 2023
Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. Libero: Bench- marking knowledge transfer for lifelong robot learning.NeurIPS, 36:44776–44791, 2023. 2, 4
2023
-
[66]
Mamba4d: Efficient 4d point cloud video understanding with disentangled spatial-temporal state space models
Jiuming Liu, Jinru Han, Lihao Liu, Angelica I Aviles-Rivero, Chaokang Jiang, Zhe Liu, and Hesheng Wang. Mamba4d: Efficient 4d point cloud video understanding with disentangled spatial-temporal state space models. InCVPR, pages 17626–17636, 2025. 3
2025
-
[67]
Difflow3d: Hierarchical diffusion models for uncertainty-aware 3d scene flow estimation.TPAMI, 2025
Jiuming Liu, Weicai Ye, Guangming Wang, Chaokang Jiang, Lei Pan, Jinru Han, Zhe Liu, Guofeng Zhang, and Hesheng Wang. Difflow3d: Hierarchical diffusion models for uncertainty-aware 3d scene flow estimation.TPAMI, 2025. 3
2025
-
[68]
Arflow: Auto-regressive optical flow estimation for arbitrary-length videos via progressive next-frame forecasting
Jiuming Liu, Mengmeng Liu, Siting Zhu, Yunpeng Zhang, Jiangtao Li, Michael Ying Yang, Francesco Nex, Hao Cheng, and Hesheng Wang. Arflow: Auto-regressive optical flow estimation for arbitrary-length videos via progressive next-frame forecasting. InICLR, 2026. 3
2026
-
[69]
Jiuming Liu, Chaojun Ni, Mengmeng Liu, Chensheng Peng, Fangjinhua Wang, Sitian Shen, Marc Pollefeys, Masayoshi Tomizuka, Ayush Tewari, and Per Ola Kristensson. Towards interactive video world modeling: Frontiers, challenges, benchmarks, and future trends.arXiv preprint arXiv:2606.01164, 2026. 2
Pith/arXiv arXiv 2026
-
[70]
Robotransfer: Controllable geometry-consistent video diffusion for manipulation policy transfer
Liu Liu, Xiaofeng Wang, Guosheng Zhao, Keyu Li, Wenkang Qin, Jiagang Zhu, Jiaxiong Qiu, Guan Huang, and Zhizhong Su. Robotransfer: Controllable geometry-consistent video diffusion for manipulation policy transfer. InCVPR, pages 1410–1420, 2026. 3 36 GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation
2026
-
[71]
Physgen: Rigid-body physics- grounded image-to-video generation
Shaowei Liu, Zhongzheng Ren, Saurabh Gupta, and Shenlong Wang. Physgen: Rigid-body physics- grounded image-to-video generation. InECCV, pages 360–378. Springer, 2024. 2
2024
-
[72]
Beyond fvd: Enhanced evaluation metrics for video generation quality, 2024
Ge Ya Luo, Gian Mario Favero, Zhi Hao Luo, Alexia Jolicoeur-Martineau, and Christopher Pal. Beyond fvd: Enhanced evaluation metrics for video generation quality, 2024. URLhttps://arxiv.org/abs/ 2410.05203. 7
Pith/arXiv arXiv 2024
-
[73]
Jindi Lv, Hao Li, Jie Li, Yifei Nie, Fankun Kong, Yang Wang, Xiaofeng Wang, Zheng Zhu, Chaojun Ni, Qiuping Deng, Hengtao Li, Jiancheng Lv, and Guan Huang. Viva: A video-generative value model for robot reinforcement learning.arXiv preprint arXiv:2604.08168, 2026. 3
Pith/arXiv arXiv 2026
-
[74]
Follow your pose: Pose-guided text-to-video generation using pose-free videos
Yue Ma, Yingqing He, Xiaodong Cun, Xintao Wang, Siran Chen, Xiu Li, and Qifeng Chen. Follow your pose: Pose-guided text-to-video generation using pose-free videos. InAAAI, volume 38, pages 4117–4125, 2024a. 3
-
[75]
Follow-your-emoji: Fine-controllable and expressive freestyle portrait animation
Yue Ma, Hongyu Liu, Hongfa Wang, Heng Pan, Yingqing He, Junkun Yuan, Ailing Zeng, Chengfei Cai, Heung-Yeung Shum, Wei Liu, et al. Follow-your-emoji: Fine-controllable and expressive freestyle portrait animation. InSIGGRAPH Asia, pages 1–12, 2024b
-
[76]
Controllable video generation: A survey.arXiv preprint arXiv:2507.16869, 2025a
Yue Ma, Kunyu Feng, Zhongyuan Hu, Xinyu Wang, Yucheng Wang, Mingzhe Zheng, Xuanhua He, Chenyang Zhu, Hongyu Liu, Yingqing He, et al. Controllable video generation: A survey.arXiv preprint arXiv:2507.16869, 2025a
-
[77]
Yue Ma, Yulong Liu, Qiyuan Zhu, Ayden Yang, Kunyu Feng, Xinhua Zhang, Zhifeng Li, Sirui Han, Chenyang Qi, and Qifeng Chen. Follow-your-motion: Video motion transfer via efficient spatial-temporal decoupled finetuning.arXiv preprint arXiv:2506.05207, 2025b
-
[78]
Yue Ma, Zexuan Yan, Hongyu Liu, Hongfa Wang, Heng Pan, Yingqing He, Junkun Yuan, Ailing Zeng, Chengfei Cai, Heung-Yeung Shum, et al. Follow-your-emoji-faster: Towards efficient, fine-controllable, and expressive freestyle portrait animation.arXiv preprint arXiv:2509.16630, 2025c
-
[79]
Follow-your-click: Open-domain regional image animation via motion prompts
Yue Ma, Yingqing He, Hongfa Wang, Andong Wang, Leqi Shen, Chenyang Qi, Jixuan Ying, Chengfei Cai, Zhifeng Li, Heung-Yeung Shum, et al. Follow-your-click: Open-domain regional image animation via motion prompts. InAAAI, volume 39, pages 6018–6026, 2025d
-
[80]
Yue Ma, Kunyu Feng, Xinhua Zhang, Hongyu Liu, David Junhao Zhang, Jinbo Xing, Yinhan Zhang, Ayden Yang, Zeyu Wang, and Qifeng Chen. Follow-your-creation: Empowering 4d creation through video inpainting.arXiv preprint arXiv:2506.04590, 2025e
-
[81]
Group editing: Edit multiple images in one go.arXiv preprint arXiv:2603.22883, 2026a
Yue Ma, Xinyu Wang, Qianli Ma, Qinghe Wang, Mingzhe Zheng, Xiangpeng Yang, Hao Li, Chongbo Zhao, Jixuan Ying, Harry Yang, et al. Group editing: Edit multiple images in one go.arXiv preprint arXiv:2603.22883, 2026a
-
[82]
Fastvmt: Eliminating redundancy in video motion transfer.arXiv preprint arXiv:2602.05551, 2026b
Yue Ma, Zhikai Wang, Tianhao Ren, Mingzhe Zheng, Hongyu Liu, Jiayi Guo, Mark Fong, Yuxuan Xue, Zixiang Zhao, Konrad Schindler, et al. Fastvmt: Eliminating redundancy in video motion transfer.arXiv preprint arXiv:2602.05551, 2026b. 3
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.