Pith. sign in

REVIEW 3 major objections 5 minor 162 references

Robot policy evaluators succeed by staying action-faithful over long horizons, not by looking more photorealistic.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 08:03 UTC pith:KFY7ODQ4

load-bearing objection Solid large-scale empirical roadmap for world models as robot policy evaluators; the 14.9% headline is on a chosen diagnostic average, not quantified closed-loop ranking ρ. the 3 major comments →

arxiv 2607.02642 v1 pith:KFY7ODQ4 submitted 2026-07-02 cs.RO

GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

classification cs.RO
keywords world modelsrobot policy evaluationWMBenchaction-conditioned video generationlong-horizon rolloutembodied AIGigaWorld-1closed-loop evaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Robot foundation models cannot be scored the way language models are scored: every checkpoint still needs slow, supervised physical rollouts. This paper argues that learned video world models can stand in for those rollouts only when they preserve the success and failure ranking of real policies, not when they merely produce pretty frames. The authors build WMBench from paired teleoperation and policy trajectories, then run a controlled study of seven world models, four action encodings, and more than 324,000 simulated rollouts against real executions. Three results follow. First, evaluator quality tracks long-horizon action fidelity far more than short-term visual realism; metrics that reward static or action-ignorant videos actively mislead ranking. Second, pretraining helps only when broad physical knowledge is kept in balance with robot-specific controllability. Third, design choices—pixel-aligned action maps, hierarchical memory with a first-frame anchor, and evaluator-focused post-training—decide whether simulated outcomes match real ones. The authors turn those rules into GigaWorld-1, trained on roughly 13,000 hours of mixed data, which raises the core evaluator-alignment average by 14.9 percent over the strongest matched general-purpose baseline.

Core claim

A world model is a reliable robot-policy evaluator only when its closed-loop rollouts stay action-faithful over long horizons and therefore reproduce the same success/failure ranking that real robots produce; short-term photorealism is secondary, and the decisive levers are balanced physical-plus-robot data, spatially aligned action control, persistent multi-scale memory, and post-training aimed at evaluator agreement rather than generic video quality.

What carries the argument

WMBench: a paired real/world-model rollout benchmark whose primary target is ranking correlation between real-world and world-model success rates (and the related ordinal WMES score), used to isolate which metrics, data mixes, and architectures actually predict real policy outcomes.

Load-bearing premise

Agreement on WMBench’s eight held-out manipulation families, still drawn from the same platforms, cameras, and task distribution, is enough to claim that a world model will rank novel policies reliably under broader initial states, embodiments, and contact-rich failures.

What would settle it

Take a set of policies whose real-robot success ranking is known on tasks outside WMBench’s eight families or on a different embodiment; if the world model’s closed-loop ranking correlation collapses while short-horizon visual metrics remain high, the central claim that long-horizon action faithfulness on this benchmark is sufficient fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Policy developers can replace a large fraction of hardware rollouts with closed-loop world-model evaluation once the model is scored by ranking agreement rather than frame beauty.
  • Metric suites that reward static or action-ignorant videos will systematically promote weak evaluators and should be dropped from evaluator leaderboards.
  • Training recipes must mix broad physical video with robot data; robot-only fine-tuning improves embodiment look but can erase the priors needed for reliable evaluation.
  • Spatially aligned control maps plus hierarchical memory become standard requirements for any video world model intended as a policy surrogate.
  • Open release of the benchmark, models, and annotation toolkit lets the community iterate on evaluator design the way language-model groups iterate on digital suites.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same long-horizon action-faithfulness test could become a filter for world models used as data engines or planners, not only as offline evaluators.
  • If VLM outcome labeling stays within a few percent of human WMES at method level, most future evaluator leaderboards can run without exhaustive human annotation.
  • Contact-rich failure modes still show optimistic bias; hybrid world models that add explicit contact or force state may be the next necessary step beyond pure video.
  • Once ranking correlation is the accepted target, sim-to-real gaps in classical simulators can be measured against the same yardstick rather than against visual fidelity alone.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies world models as surrogate evaluators of robot policies, arguing that real-robot evaluation is the bottleneck for embodied foundation models. It introduces WMBench (paired teleoperation and policy-rollout data over eight manipulation task families), analyzes seven video world models, four action encodings, and 324k+ annotated rollouts (including CVPR 2026 challenge submissions), and reports three design insights: long-horizon action-faithful consistency dominates short-term visual realism; pretraining gains require balancing general physical priors with robot controllability; and action interface, memory, and evaluator-oriented post-training strongly affect real-world alignment. These are instantiated in GigaWorld-1 (Wan-based, ~13k hours multi-source data, pixel-aligned EE/ray control, hierarchical memory, progressive training), which improves a six-metric average by 14.9% over Wan 2.2 5B under matched post-training (Table 9). Code, models, and data are released.

Significance. If the design claims hold, the work is a substantial contribution to scalable robot policy evaluation: it reframes world models as external policy evaluators rather than only data engines or planners, provides a large paired real/sim benchmark and metric analysis (Figs. 4–5, WMES), and ships a concrete open roadmap (data mixture, spatially aligned control, memory, distillation) with reproducible artifacts. The controlled ablations (Tables 2–4, 9) and community-scale annotation are genuine strengths relative to prior proof-of-concept evaluator papers. The practical value depends on whether closed-loop ranking fidelity—not only diagnostic video metrics—is demonstrated at the same scale as the headline gains.

major comments (3)
  1. [Sec. 3 Eq. (4); Sec. 6.5.1 Table 9; Sec. 6.5.5 Figs. 16–17] Sec. 3, Eq. (4) defines the primary evaluator target as ranking/success agreement ρ = Corr(S_real(π), S_wm(π)) across policies. The abstract, intro, and Sec. 6.5.1 instead headline a 14.9% gain of GigaWorld-1-Plus over Wan 2.2 5B on a paper-chosen six-metric average (Aesthetic, Image, JEPA, Semantic, Subject, Trajectory; Table 9: 0.6834 vs 0.5948). Those diagnostics are justified by correlation with human WMES on challenge rollouts (Findings 1–2, Figs. 4–5), but they are not ρ. Closed-loop evidence in Sec. 6.5.5 is limited to task-level success-rate scatter and Gen−Real bias bars on four tasks/subtasks (Figs. 16–17), without a quantified ρ (or Kendall/Spearman ranking) over a multi-checkpoint policy set for GigaWorld-1 vs baselines. Please report ρ (with CIs) under the same closed-loop protocol for the main models, or reframe the 14.9% claim so it is not presented as the primary evaluato
  2. [Sec. 6.5.5; Fig. 17; Finding 3 / Table 4] Sec. 6.5.5 and Fig. 17 acknowledge residual optimistic bias on contact-sensitive failures (e.g., pour/press subtasks). Because the paper’s own Finding 3 and Table 4 argue that long-horizon action-faithful consistency—not short-horizon visual quality—dominates evaluator reliability, the closed-loop calibration gap is load-bearing. Either quantify how often GigaWorld-1 flips real success/failure relative to baselines (confusion matrices or per-subtask agreement rates on the full WMBench closed-loop set), or temper claims that the model is “specially optimized for policy evaluation” until failure-mode calibration is measured at the same scale as the diagnostic average.
  3. [Sec. 4.3; Findings 4–5; Table 1; Table 9] WMES and several diagnostic metrics (Perspectivity, Instruction Following, Interaction Quality, Semantic Alignment) rely on VLM judges (Sec. 4.3; Finding 4–5; Table 1). The LoRA VLM is trained to predict human WMES with score-token weight 8.0, then used for scalable outcome assessment. Metric selection for Table 9 is guided by correlation with that same WMES family. This is not fatal circularity—human WMES and real success labels remain external—but it creates mild self-reinforcement risk if models are tuned toward VLM-preferred appearance/geometry. Please report (i) human-only vs VLM-only ranking of the main models on a held-out subset, and (ii) sensitivity of the six-metric average and any ρ to excluding VLM-judged diagnostics.
minor comments (5)
  1. [Fig. 1; Abstract; Sec. 4.1; Table 6] Fig. 1 and abstract claim “324,000+ analyzed rollouts” and “12K+ hours training data”; Table 6 totals ~12,980 hours. Align the rounded figures and clarify whether 324k counts segments or full closed-loop episodes (Sec. 4.1 says segments chained into episodes of 20–30 segments).
  2. [Table 3; Finding 9] Table 3 ranks control interfaces on Trajectory Accuracy and motion metrics for Wan 2.1 1.3B only. A short note on whether channel-concat remains best for the 5B backbone (GigaWorld-1-Plus) would strengthen Finding 9.
  3. [Sec. 4.1; Sec. 6.5.4; Sec. 7] Sec. 4.1 train/test split is episode-disjoint within the same eight task families and platforms. Sec. 7 correctly flags limited coverage of mobile/dexterous/safety-critical settings; a brief explicit statement in Sec. 6.5 that OOD claims (Fig. 15) are appearance/content shifts, not embodiment or policy-family shifts, would prevent over-reading.
  4. [Fig. 2; Fig. 3] Typo in Fig. 3 caption/step labels: “Train Wodel Model” should be “Train World Model.” Several figure panels (e.g., Fig. 2 Chinese annotation fragment in the source) should be cleaned for the camera-ready version.
  5. [Finding 6 footnote; Sec. 7] Cosmos-3 is mentioned as planned but unavailable (footnote in Finding 6). Either update the comparison if multiview access is obtained, or move the note to a single limitations paragraph to avoid dangling promises.

Circularity Check

1 steps flagged

No load-bearing circular derivation: claims rest on external real-robot pairings and held-out metrics; only mild metric-suite selection via human WMES.

specific steps
  1. other [Sec. 5.1 Findings 1–2; Sec. 6.5.1 / Table 9; Eq. 4 vs. reported AVG]
    "evaluator quality is dominated by long-horizon, action-faithful rollout consistency rather than short-term visual realism... GigaWorld-1-Plus improves the average score by ... 14.9% over Wan 2.2 5B... we retain six core evaluator-relevant metric(Aesthetic Quality, Image Quality, JEPA Similarity, Semantic Alignment, Subject Consistency, and Trajectory Accuracy)... Visual Fidelity has the highest correlation (ρ=0.78)... Subject Consistency (ρ=0.88) and Perspectivity (ρ=0.86) are the strongest predictors"

    The paper’s stated primary target is ρ=Corr(S_real,S_wm) (Eq. 4), but the headline 14.9% is the mean of six automatic metrics pre-selected because they correlate with human WMES on the same challenge ecosystem. That makes the reported “evaluator-alignment” gain partly an improvement on a human-rubric-derived metric suite rather than a direct closed-loop ranking result. Mild self-reinforcement of the evaluation criterion, not a by-construction identity between fit and prediction.

full rationale

This is an empirical systems paper, not a first-principles derivation. The primary evaluator target (Eq. 4) is ranking/success agreement between world-model and real-robot outcomes—an external quantity. WMBench uses episode-disjoint held-out trajectories with real teleoperation and policy rollouts; the 324k annotated segments and human WMES are independent labels, not fitted parameters renamed as predictions. Ablations (control interfaces, memory, data mixtures) and the Table 9 comparison (matched post-training of multiple backbones) measure independently computed diagnostics (JEPA, NDTW trajectory accuracy, etc.). Self-citations (GigaBrain, GigaWorld-0, etc.) supply data sources and related work, not uniqueness theorems that force the result. The only mild circularity risk is operational: Findings 1–2 select the six-metric AVG reported as “evaluator-alignment” because those metrics correlate with human WMES on challenge rollouts, and a LoRA VLM is trained to predict that same WMES for scalable closed-loop scoring—so the reported 14.9% is alignment with a paper-chosen human-rubric proxy, not a direct measurement of ρ. That is metric-selection self-reinforcement, not definitional circularity or a fitted input called a prediction; the model scores are not forced by construction. Score 1 reflects that minor operational loop without elevating it to a load-bearing circular chain.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 3 invented entities

This is empirical systems work, not a theorem paper. Load-bearing content is mostly domain assumptions about what constitutes a good evaluator (real–sim ranking agreement), engineering choices (action maps, memory, data mix), and many training/filter thresholds. No new physical entity is postulated; the main invented constructs are the benchmark, WMES rubric, and model stack.

free parameters (6)
  • Six-metric evaluator average weights (equal mean of Aesthetic, Image, JEPA, Semantic, Subject, Trajectory)
    The headline 0.6834 / 14.9% claim depends on this hand-chosen equal average after excluding appearance-stability metrics; different metric sets would change ranking.
  • VLM score-token loss weight 8.0 (format 1.0, rationale min 0.05)
    Directly shapes the automated WMES proxy used for scalable labeling and method comparison (Sec. 5.1).
  • Data mixture composition (PhysData / AgiBot / GigaData hours and filters)
    Ablations show large swings in JEPA/Trajectory vs appearance metrics; the recommended GigaData+PhysData mix is an empirical fit to their benchmark domain.
  • Video quality / motion filter thresholds (τ_img, τ_aes, τ_jump, τ_static, τ_motion, τ_jerk)
    These gates define the training corpus; values are design choices that affect what the world model learns.
  • LoRA ranks/alphas and stage learning rates (e.g., Stage1 LR 5e-5, Stage2 1e-4, ranks 128/256)
    Standard but claim-relevant training knobs for GigaWorld-1 performance (Tables 7–8).
  • Hierarchical memory partition (short/mid/long + first-frame anchor) and generation window length
    Architectural hyperparameters that drive the long-horizon PSNR/FID/FVD gains in Table 4.
axioms (5)
  • domain assumption Agreement between world-model and real-world policy success/ranking is the primary definition of evaluator quality.
    Stated in Sec. 3 (Eq. 4) and used throughout WMBench; alternative targets (e.g., calibrated risk, counterfactual returns) are not the main objective.
  • domain assumption Video diffusion backbones with action conditioning can capture decision-relevant physical dynamics for closed-loop manipulation evaluation.
    Background premise of the entire program (Intro, Related Work); failures are treated as design problems, not refutations of the premise.
  • domain assumption Human WMES ordinal labels (0–3) and LoRA-tuned VLM scores are valid ground truth for ranking world models as evaluators.
    Sec. 4.1 and Finding 5; adjacent agreement is high, but the construct still embeds human notions of fidelity plus outcome.
  • ad hoc to paper Episode-disjoint train/test splits within the same eight task families suffice to measure generalization for evaluator claims.
    Sec. 4.1 protocol; stronger robot/task OOD is only partially probed in Sec. 6.5.4.
  • standard math Standard flow-matching / diffusion training and autoregressive windowing preserve enough dynamics for long-horizon evaluation when memory and control are added.
    Training objectives in Sec. 6.3 (flow-matching and AR denoising losses) are taken as given machinery.
invented entities (3)
  • WMBench independent evidence
    purpose: Paired real-robot and world-model rollout benchmark for controlled evaluator comparison.
    New dataset/protocol construct; independent use depends on public release and external adoption.
  • WMES (World Model as Evaluator Score) no independent evidence
    purpose: Four-level human ordinal rubric combining outcome correctness and visual/physical fidelity.
    Paper-defined ground-truth construct used to select metrics and train the VLM judge; not a physical entity.
  • GigaWorld-1 (Nano/Plus) independent evidence
    purpose: Evaluator-oriented world model instantiating the design roadmap.
    Model artifact; evidence is internal benchmark gains plus promised open weights.

pith-pipeline@v1.1.0-grok45 · 44561 in / 4035 out tokens · 55193 ms · 2026-07-12T08:03:22.970962+00:00 · methodology

0 comments
read the original abstract

Evaluating embodied robot foundation models remains a critical bottleneck; unlike large language models efficiently assessed via digital benchmarks, robotic policies require slow, costly real-world rollouts limited by hardware and human supervision, which has driven interest in world models as surrogate policy evaluators, yet the key properties that make a world model reliable for policy assessment remain poorly understood. This work presents a systematic study of world models for robotic policy evaluation and introduces WMBench, a benchmark constructed from real-robot teleoperation data and matched policy rollouts covering diverse manipulation tasks to enable controlled comparisons across model families, action encodings, rollout horizons, and evaluation metrics. Using WMBench, we analyze 7 video world models, 4 action representation schemes, and over 324,000 simulated policy rollouts paired with real robot executions, further enriching our analysis with large-scale community submissions from the CVPR 2026 GigaBrain Challenge, curated synthetic trajectories, and a training videos spanning more than 12,000 hours. Our experiments deliver three core insights: evaluator quality is dominated by long-horizon, action-faithful rollout consistency rather than short-term visual realism; pretraining gains stem not only from data scale but from balancing general world knowledge with robot-specific controllability; and architectural choices including action encoding, memory design, and evaluator-focused post-training strongly determine alignment with real-world robot behavior. Drawing on these results, we derive a practical design roadmap and realize it in \textit{GigaWorld-1}, a world model specially optimized for policy evaluation, and we fully release our code, models, datasets, and toolkits to advance scalable evaluation research for embodied foundation models.

Figures

Figures reproduced from arXiv: 2607.02642 by Angyuan Ma, Bohan Li, Boyuan Wang, Chaojun Ni, GigaWorld Team, Guan Huang, Guo Li, Guosheng Zhao, Hao Li, Hengtao Li, Jingyu Liu, Jiwen Lu, Qiuping Deng, Tingdong Yu, Xiaofeng Wang, Xiaoyu Tian, Xinyu Zhou, Xinze Chen, Xiuwei Xu, Xuancheng Xu, Yang Wang, Yifan Chang, Yukun Zhou, Yun Ye, Zhanqian Wu, Zheng Zhu, Zhenyu Wu.

Figure 1
Figure 1. Figure 1: This paper analyzes 324,000 world-model-simulated rollouts, 7 video world models, 4 action [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: World model as policy evaluator framework. A world model serves as a policy evaluator by iteratively receiving policy actions and predicting future observations. Reliable evaluation requires not only visual quality, but also action-faithful rollout and agreement with real-world policy outcomes. and difficult to scale comprehensively under broad distribution shifts. To address this scalability bottleneck, t… view at source ↗
Figure 3
Figure 3. Figure 3: WMBench evaluation pipeline. The four-step protocol includes (1) collecting real-world policy rollouts, (2) training world models on a strict split, (3) executing closed-loop policy rollouts inside the learned world model, and (4) assessing metrics and outcomes to measure alignment with real-world conclusions. tor [49]. Aesthetic Quality measures visual appeal, lighting, and color composition using the LAI… view at source ↗
Figure 4
Figure 4. Figure 4: Metric-group correlation with WMES. Metrics submitted to the WMBench are grouped into visual fidelity, geometry, semantics, dynamics, interaction, and appearance stability categories. Visual fidelity and geometry are the strongest group-level predictors, while appearance stability is negatively correlated with WMES. individual-metric level, Subject Consistency (𝜌 = 0.88) and Perspectivity (𝜌 = 0.86) are th… view at source ↗
Figure 5
Figure 5. Figure 5: Pearson correlation matrix over all submitted metrics. The full metric-level heatmap shows that Subject Consistency, Perspectivity, JEPA Similarity, Instruction Following, Image Quality, and Aesthetic Quality correlate strongly with WMES, whereas appearance-stability metrics and Interaction Quality are much less reliable. VL-8B-Instruct using LoRA with structured supervision that couples the overall WMES s… view at source ↗
Figure 6
Figure 6. Figure 6: VLM-assisted Rollout Evaluator. Given a three-view rollout video and a task-specific evaluation prompt, the LoRA-tuned Qwen3-VL evaluator predicts a WMES score and produces evidence-grounded rationales with structured aspect-level assessments of overall video quality, instruction following, and physical adherence. Acc. ↑ Adj. Acc. ↑ Large Err. ↓ MAE ↓ RMSE ↓ QWK ↑ Spearman ↑ Kendall 𝜏𝑏 ↑ W-F1 ↑ |Bias| ↓ 0.… view at source ↗
Figure 7
Figure 7. Figure 7: Data construction pipeline. Multi-source data is filtered, balanced, and automatically annotated before being incorporated into the world-model training corpus. 6.1. Data Sources and Data Curation 6.1.1 Data Composition The success of large-scale embodied foundation models relies heavily on the quality, diversity, and scale of training data. The ablation results in Sec. 5 show that broad physical-video pri… view at source ↗
Figure 8
Figure 8. Figure 8: Overall architecture of GigaWorld-1. The model is built as an autoregressive diffusion-transformer world generator with parameter-efficient LoRA adaptation. Historical frames are encoded through memory patchification, future noisy latents are encoded through patchification, and structured controls such as actions, depth, semantic maps, and captions are injected as temporally aligned conditions. Frozen VAE … view at source ↗
Figure 9
Figure 9. Figure 9: Prompt transition via spherical linear interpolation. Instead of abruptly switching between prompt embeddings, GigaWorld-1 samples intermediate text conditions along the spherical path between two semantic endpoints. These interpolated embeddings are injected into successive autoregressive windows, producing smoother changes in scene appearance, motion pattern, and task phase. 𝜃 = arccos (︂ e ⊤ 1 e2 ‖e1‖‖e… view at source ↗
Figure 10
Figure 10. Figure 10: Training pipeline of GigaWorld-1. The model is first adapted into a robot world foundation model, converted into an autoregressive world generator, and finally compressed through optional ODE warm start and required DMD2 distillation for few-step rollout. Dashed branches denote optional modules that can be skipped. Stage 1: robot world foundation model. We first initialize from a pretrained video backbone… view at source ↗
Figure 11
Figure 11. Figure 11: Model architecture comparison. Left: mean score across six evaluation metrics. Right: radar plot of individual metric scores. GigaWorld-1-Nano achieves the second-best overall score and performs particularly well in JEPA Similarity, Semantic Alignment, Subject Consistency, and Trajectory Accuracy [PITH_FULL_IMAGE:figures/full_fig_p026_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Long-horizon rollout dynamics. PSNR (↑) and FID (↓) are measured over successive 10-frame rollout chunks, showing how reconstruction fidelity and perceptual quality evolve with rollout length. identified in Question I: GigaWorld-1-Plus achieves the best JEPA Similarity (0.9337), Semantic Alignment (0.8926), and Trajectory Accuracy (0.3561), while matching the best Subject Consistency score (0.8883). 6.5.2… view at source ↗
Figure 13
Figure 13. Figure 13: Long-horizon model comparison. GigaWorld-1 is compared with general video generation and world-model baselines under the same rollout evaluation protocol. inconsistent scene layouts. While adding memory stabilizes the workspace, abrupt prompt changes can cause the model to over-condition on previous subtasks. Combining memory with Spherical Linear Interpolation (SLERP) resolves this trade-off, enabling sm… view at source ↗
Figure 14
Figure 14. Figure 14: Effect of memory and prompt interpolation (left arm moves to a predefined observation pose, while the right arm places a towel into the blue box). Without memory, the rollout suffers from background jumps and inconsistent scene layout. Adding history memory stabilizes the workspace, but a fixed or abruptly switched prompt can make the model over-conditioned on the previous subtask during stage transitions… view at source ↗
Figure 15
Figure 15. Figure 15: OOD generalization cases. From top to bottom, the rows evaluate generalization to object color and container appearance, object content changes, background and table-surface changes, and action-outcome variation covering both successful and failed executions. What matters is the ability to preserve broad world knowledge, remain controllable under action input, and sustain long-horizon consistency while ke… view at source ↗
Figure 16
Figure 16. Figure 16: Task-level success-rate alignment. Real-robot success rates are compared with generated success rates under closed-loop policy rollout. The gray dashed line indicates perfect agreement. GigaWorld-1 has a fitted line closer to the diagonal than the challenge baselines, indicating better calibration of task difficulty across the evaluated tasks and subtasks. models; other structured state-space or hybrid 3D… view at source ↗
Figure 17
Figure 17. Figure 17: Success-rate bias across world models and subtasks. Bars show Gen − Real success-rate differences: green indicates overestimation of real-world success, red indicates underestimation, and values near zero indicate closer agreement. 31 [PITH_FULL_IMAGE:figures/full_fig_p031_17.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

162 extracted references · 74 linked inside Pith

  1. [1]

    Cosmos 3: Omnimodal world models for physical ai.arXiv preprint arXiv:2606.02800, 2026

    Niket Agarwal, Arslan Ali, Jon Allen, Martin Antolini, Adeline Aubame, Alisson Azzolini, Junjie Bai, Maciej Bala, Yogesh Balaji, Josh Bapst, et al. Cosmos 3: Omnimodal world models for physical ai.arXiv preprint arXiv:2606.02800, 2026. 2, 3, 4, 11, 30

  2. [2]

    Cosmos-transfer1: Conditional world generation with adaptive multimodal control.arXiv preprint arXiv:2503.14492, 2025

    Hassan Abu Alhaija, Jose Alvarez, Maciej Bala, Tiffany Cai, Tianshi Cao, Liz Cha, Joshua Chen, Mike Chen, Francesco Ferroni, Sanja Fidler, et al. Cosmos-transfer1: Conditional world generation with adaptive multimodal control.arXiv preprint arXiv:2503.14492, 2025. 3

  3. [3]

    World simulation with video foundation models for physical ai.arXiv preprint arXiv:2511.00062, 2025

    Arslan Ali, Junjie Bai, Maciej Bala, Yogesh Balaji, Aaron Blakeman, Tiffany Cai, Jiaxin Cao, Tianshi Cao, Elizabeth Cha, Yu-Wei Chao, et al. World simulation with video foundation models for physical ai.arXiv preprint arXiv:2511.00062, 2025. 11

  4. [4]

    Roboarena: Distributed real-world evaluation of generalist robot policies.arXiv preprint arXiv:2506.18123, 2025

    Pranav Atreya, Karl Pertsch, Tony Lee, Moo Jin Kim, Arhan Jain, Artur Kuramshin, Clemens Eppner, Cyrus Neary, Edward Hu, Fabio Ramos, et al. Roboarena: Distributed real-world evaluation of generalist robot policies.arXiv preprint arXiv:2506.18123, 2025. 3

  5. [5]

    Hyperreel: High-fidelity 6-dof video with ray-conditioned sampling

    Benjamin Attal, Jia-Bin Huang, Christian Richardt, Michael Zollhoefer, Johannes Kopf, Matthew O’Toole, and Changil Kim. Hyperreel: High-fidelity 6-dof video with ray-conditioned sampling. InCVPR, pages 16610–16620, 2023. 19

  6. [7]

    Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025

    Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junyang Lin, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan ...

  7. [9]

    Qwen2.5-vl technical report.arXiv preprint arXiv:2502.13923,

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-vl technical report.ar...

  8. [10]

    7 32 GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

  9. [11]

    Navigation world models

    Amir Bar, Gaoyue Zhou, Danny Tran, Trevor Darrell, and Yann LeCun. Navigation world models. In CVPR, pages 15791–15801, 2025. 2

  10. [12]

    V-JEPA: Latent video prediction for visual representation learning, 2024

    Adrien Bardes, Quentin Garrido, Jean Ponce, Xinlei Chen, Michael Rabbat, Yann LeCun, Mido Assran, and Nicolas Ballas. V-JEPA: Latent video prediction for visual representation learning, 2024. URL https://openreview.net/forum?id=WFYbBOEOtv. 7

  11. [13]

    Motus: A unified latent action world model.arXiv preprint arXiv:2512.13030, 2025

    Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, Shuhe Huang, Haitian Liu, Ruowen Zhao, Yao Feng, Chendong Xiang, Yinze Rong, et al. Motus: A unified latent action world model.arXiv preprint arXiv:2512.13030, 2025. 2, 3

  12. [14]

    Gr00t n1: An open foundation model for generalist humanoid robots.arXiv preprint arXiv:2503.14734, 2025

    Johan Bjorck, Fernando Castañeda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, et al. Gr00t n1: An open foundation model for generalist humanoid robots.arXiv preprint arXiv:2503.14734, 2025

  13. [15]

    pi0: A vision-language-action flow model for general robot control.arXiv preprint arXiv:2410.24164, 2024

    Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, et al. pi0: A vision-language-action flow model for general robot control.arXiv preprint arXiv:2410.24164, 2024. 2

  14. [16]

    Stable video diffusion: Scaling latent video diffusion models to large datasets.arXiv preprint arXiv:2311.15127, 2023

    Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets.arXiv preprint arXiv:2311.15127, 2023. 2, 11

  15. [17]

    Tiny autoencoder for stable diffusion.https://github.com/madebyollin/taesd,

    Ollin Boer Bohan. Tiny autoencoder for stable diffusion.https://github.com/madebyollin/taesd,

  16. [18]

    GitHub repository. 24

  17. [19]

    Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems

    Qingwen Bu, Jisong Cai, Li Chen, Xiuqi Cui, Yan Ding, Siyuan Feng, Xindong He, Xu Huang, et al. Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems. In IROS, 2025. 15

  18. [20]

    Sam 3: Segment anything with concepts, 2025

    Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Didac Suris, Chai- tanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, Jie Lei, Tengyu Ma, Baishan Guo, Arpit Kalla, Markus Marks, Joseph Greer, Meng Wang, Peize Sun, Roman Rädle, Triantafyllos Afouras, Effrosyni Mavroudi, Katherine Xu, Tsung-Han Wu, Yu Zhou, Lilian...

  19. [21]

    Emerging properties in self-supervised vision transformers, 2021

    Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers, 2021. URLhttps://arxiv.org/ abs/2104.14294. 7

  20. [22]

    Worldvla: Towards autoregressive action world model.arXiv preprint arXiv:2506.21539, 2025

    Jun Cen, Chaohui Yu, Hangjie Yuan, Yuming Jiang, Siteng Huang, Jiayan Guo, Xin Li, Yibing Song, Hao Luo, Fan Wang, et al. Worldvla: Towards autoregressive action world model.arXiv preprint arXiv:2506.21539, 2025. 2

  21. [23]

    Reflect-r1: Evidence-driven reflection for self-correction in long video understanding, 2026

    Shuimu Chen, Yuteng Chen, Yuanshen Guan, Zebang Cheng, Zeyu Zhang, Shengqian Qin, Bin Xia, Jiaran Li, Wenming Yang, and Fei Ma. Reflect-r1: Evidence-driven reflection for self-correction in long video understanding, 2026. URLhttps://arxiv.org/abs/2606.27922. 17

  22. [24]

    Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation.arXiv preprint arXiv:2506.18088,

    Tianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai, Yibin Liu, Qiwei Liang, Zixuan Li, Xianliang Lin, Yiheng Ge, Zhenyu Gu, et al. Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation.arXiv preprint arXiv:2506.18088,

  23. [25]

    2, 4 33 GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

  24. [26]

    Abot-physworld: Interactive world foundation model for robotic manipulation with physics alignment.arXiv preprint arXiv:2603.23376, 2026

    Yuzhi Chen, Ronghan Chen, Dongjie Huo, Yandan Yang, Dekang Qi, Haoyun Liu, Tong Lin, Shuang Zeng, Junjin Xiao, Xinyuan Chang, et al. Abot-physworld: Interactive world foundation model for robotic manipulation with physics alignment.arXiv preprint arXiv:2603.23376, 2026. 2

  25. [27]

    Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021. 2

  26. [28]

    Introducing veo 3, our video generation model with expanded creative controls – including native audio and extended videos

    Google DeepMind. Introducing veo 3, our video generation model with expanded creative controls – including native audio and extended videos. [Online], 2025. URLhttps://deepmind.google/models/ veo/. 3

  27. [29]

    Emma: Generalizingreal-worldrobotmanipulationviagenerative visual transfer.arXiv preprint arXiv:2509.22407, 2025

    Zhehao Dong, Xiaofeng Wang, Zheng Zhu, Yirui Wang, Yang Wang, Yukun Zhou, Boyuan Wang, Chaojun Ni, RunqiOuyang, WenkangQin, etal. Emma: Generalizingreal-worldrobotmanipulationviagenerative visual transfer.arXiv preprint arXiv:2509.22407, 2025. 3

  28. [30]

    Aim: Intent-aware unified world action modeling with spatial value maps.arXiv preprint arXiv:2604.11135, 2026

    Liaoyuan Fan, Zetian Xu, Chen Cao, Wenyao Zhang, Mingqi Yuan, and Jiayu Chen. Aim: Intent-aware unified world action modeling with spatial value maps.arXiv preprint arXiv:2604.11135, 2026. 3

  29. [31]

    Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

    Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. InCVPR, pages 24108–24118, 2025. 2

  30. [32]

    Learning video generation for robotic manipulation with collaborative trajectory control.arXiv preprint arXiv:2506.01943, 2025

    Xiao Fu, Xintao Wang, Xian Liu, Jianhong Bai, Runsen Xu, Pengfei Wan, Di Zhang, and Dahua Lin. Learning video generation for robotic manipulation with collaborative trajectory control.arXiv preprint arXiv:2506.01943, 2025. 19

  31. [33]

    Seedance 1.0: Exploring the boundaries of video generation models.arXiv preprint arXiv:2506.09113, 2025

    Yu Gao, Haoyuan Guo, Tuyen Hoang, Weilin Huang, Lu Jiang, Fangyuan Kong, Huixia Li, Jiashi Li, Liang Li, Xiaojie Li, et al. Seedance 1.0: Exploring the boundaries of video generation models.arXiv preprint arXiv:2506.09113, 2025. 3

  32. [34]

    Genesis: Multimodal driving scene generation with spatio-temporal and cross-modal consistency.NeurIPS, 38:60431–60455, 2026

    Xiangyu Guo, Zhanqian Wu, Kaixin Xiong, Ziyang Xu, Lijun Zhou, Gangwei Xu, Shaoqing Xu, Haiyang Sun, Bing Wang, Guang Chen, et al. Genesis: Multimodal driving scene generation with spatio-temporal and cross-modal consistency.NeurIPS, 38:60431–60455, 2026. 3

  33. [35]

    Ctrl-world: A controllable generative world model for robot manipulation.arXiv preprint arXiv:2510.10125, 2025

    Yanjiang Guo, Lucy Xiaoyang Shi, Jianyu Chen, and Chelsea Finn. Ctrl-world: A controllable generative world model for robot manipulation.arXiv preprint arXiv:2510.10125, 2025. 2, 3, 4

  34. [36]

    Ltx-video: Realtime video latent diffusion.arXiv preprint arXiv:2501.00103, 2024

    Yoav HaCohen, Nisan Chiprut, Benny Brazowski, Daniel Shalem, Dudu Moshe, Eitan Richardson, Eran Levin, Guy Shiran, Nir Zabari, Ori Gordon, et al. Ltx-video: Realtime video latent diffusion.arXiv preprint arXiv:2501.00103, 2024. 2, 11

  35. [37]

    Pre-trained video generative models as world simulators

    Haoran He, Yang Zhang, Liang Lin, Zhongwen Xu, and Ling Pan. Pre-trained video generative models as world simulators. InAAAI, volume 40, pages 4645–4653, 2026. 2

  36. [38]

    Matrix-game 2.0: An open-source real-time and streaming interactive world model.arXiv preprint arXiv:2508.13009, 2025

    Xianglong He, Chunli Peng, Zexiang Liu, Boyang Wang, Yifan Zhang, Qi Cui, Fei Kang, Biao Jiang, Mengyin An, Yangyang Ren, et al. Matrix-game 2.0: An open-source real-time and streaming interactive world model.arXiv preprint arXiv:2508.13009, 2025. 3

  37. [39]

    Skip: Sparse keyframe interpolation paradigm for efficient embodied world models.arXiv preprint arXiv:2606.00664, 2026

    Ziheng He, Yixiang Chen, Ning Yang, Zhanqian Wu, Qisen Ma, Yuan Xu, Jiabing Yang, Peiyan Li, Xiangnan Wu, Xiaofeng Wang, et al. Skip: Sparse keyframe interpolation paradigm for efficient embodied world models.arXiv preprint arXiv:2606.00664, 2026. 3

  38. [40]

    Measuring massive multitask language understanding

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. InICLR, 2020. 2 34 GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

  39. [41]

    Gans trained by a two time-scale update rule converge to a local nash equilibrium.NeurIPS, 2017

    Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium.NeurIPS, 2017. 8

  40. [42]

    Cogvideo: Large-scale pretraining for text-to-video generation via transformers

    Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. InICLR, 2023. 3

  41. [43]

    Ruler: What’s the real context size of your long-context language models?arXiv preprint arXiv:2404.06654, 2024

    Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg. Ruler: What’s the real context size of your long-context language models?arXiv preprint arXiv:2404.06654, 2024. 2

  42. [44]

    Lora: Low-rank adaptation of large language models

    Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. Lora: Low-rank adaptation of large language models. InICLR, 2021. 18

  43. [45]

    Self forcing: Bridging the train-test gap in autoregressive video diffusion.NeurIPS, 38:167283–167308, 2026

    Xun Huang, Zhengqi Li, Guande He, Mingyuan Zhou, and Eli Shechtman. Self forcing: Bridging the train-test gap in autoregressive video diffusion.NeurIPS, 38:167283–167308, 2026. 3

  44. [46]

    Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, et al.𝜋0.5: a vision-language-action model with open-world generalization.arXiv preprint arXiv:2504.16054, 2025. 2

  45. [47]

    Rlbench: The robot learning benchmark & learning environment.IEEE Robotics and Automation Letters, 5(2):3019–3026, 2020

    Stephen James, Zicong Ma, David Rovick Arrojo, and Andrew J Davison. Rlbench: The robot learning benchmark & learning environment.IEEE Robotics and Automation Letters, 5(2):3019–3026, 2020. 2

  46. [48]

    Galaxea open-world dataset and g0 dual-system vla model.arXiv preprint arXiv:2509.00576, 2025

    Tao Jiang, Tianyuan Yuan, Yicheng Liu, Chenhao Lu, Jianning Cui, Xiao Liu, Shuiqi Cheng, Jiyang Gao, Huazhe Xu, and Hang Zhao. Galaxea open-world dataset and g0 dual-system vla model.arXiv preprint arXiv:2509.00576, 2025. 2, 15

  47. [49]

    Enerverse-ac: Envisioning embodied environments with action condition.arXiv preprint arXiv:2505.09723, 2025

    Yuxin Jiang, Shengcong Chen, Siyuan Huang, Liliang Chen, Pengfei Zhou, Yue Liao, Xindong He, Chiming Liu, Hongsheng Li, Maoqing Yao, et al. Enerverse-ac: Envisioning embodied environments with action condition.arXiv preprint arXiv:2505.09723, 2025. 2

  48. [50]

    Swe-bench: Can language models resolve real-world github issues? InICLR, volume 2024, pages 54107–54157, 2024

    Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. Swe-bench: Can language models resolve real-world github issues? InICLR, volume 2024, pages 54107–54157, 2024. 2

  49. [51]

    How far is video generation from world model: A physical law perspective.arXiv preprint arXiv:2411.02385,

    Bingyi Kang, Yang Yue, Rui Lu, Zhijie Lin, Yang Zhao, Kaixin Wang, Gao Huang, and Jiashi Feng. How far is video generation from world model: A physical law perspective.arXiv preprint arXiv:2411.02385,

  50. [52]

    Musiq: Multi-scale image quality transformer

    Junjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar, and Feng Yang. Musiq: Multi-scale image quality transformer. InICCV, pages 5148–5157, October 2021. 7

  51. [53]

    Openvla: Anopen-sourcevision-language-action model.arXiv preprint arXiv:2406.09246, 2024

    Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov,EthanFoster,GraceLam,PannagSanketi,etal. Openvla: Anopen-sourcevision-language-action model.arXiv preprint arXiv:2406.09246, 2024. 2

  52. [54]

    Cosmos policy: Fine-tuning video models for visuomotor control and planning.arXiv preprint arXiv:2601.16163, 2026

    Moo Jin Kim, Yihuai Gao, Tsung-Yi Lin, Yen-Chen Lin, Yunhao Ge, Grace Lam, Percy Liang, Shuran Song, Ming-Yu Liu, Chelsea Finn, et al. Cosmos policy: Fine-tuning video models for visuomotor control and planning.arXiv preprint arXiv:2601.16163, 2026. 2, 3

  53. [55]

    Aesthetic predictor, 2022

    LAION-AI. Aesthetic predictor, 2022. URLhttps://github.com/LAION-AI/aesthetic-predictor. Accessed: 2024. 7

  54. [56]

    Vag: Dual-stream video-action generation for embodied data synthesis.arXiv preprint arXiv:2604.09330, 2026

    Xiaolei Lang, Yang Wang, Yukun Zhou, Chaojun Ni, Kerui Li, Jiagang Zhu, Tianze Liu, Jiajun Lv, Xingxing Zuo, Yun Ye, et al. Vag: Dual-stream video-action generation for embodied data synthesis.arXiv preprint arXiv:2604.09330, 2026. 2 35 GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

  55. [57]

    Robomemarena: A comprehensive and challenging robotic memory benchmark.arXiv preprint arXiv:2605.10921, 2026

    HuashuoLei,WenxuanSong,HuaruiZhang,JieyuanPei,JiayiChen,HaodongYan,HanZhao,Pengxiang Ding, Zhipeng Zhang, Lida Huang, et al. Robomemarena: A comprehensive and challenging robotic memory benchmark.arXiv preprint arXiv:2605.10921, 2026. 4

  56. [58]

    Mimicdreamer: Aligning human and robot demonstrations for scalable vla training.arXiv preprint arXiv:2509.22199, 2025

    HaoyunLi, IvanZhang, RunqiOuyang, XiaofengWang, ZhengZhu, ZhiqinYang, ZhentaoZhang, Boyuan Wang, Chaojun Ni, Wenkang Qin, et al. Mimicdreamer: Aligning human and robot demonstrations for scalable vla training.arXiv preprint arXiv:2509.22199, 2025. 3

  57. [59]

    Stableidm: Stabilizing inverse dynamics model against manipulator truncation via spatio-temporal refinement.arXiv preprint arXiv:2604.17887, 2026

    Kerui Li, Zhe Jing, Xiaofeng Wang, Zheng Zhu, Yukun Zhou, Guan Huang, Dongze Li, Qingkai Yang, and Huaibo Huang. Stableidm: Stabilizing inverse dynamics model against manipulator truncation via spatio-temporal refinement.arXiv preprint arXiv:2604.17887, 2026. 3

  58. [60]

    Evaluatingreal-worldrobotmanipulationpoliciesinsimulation

    Xuanlin Li, Kyle Hsu, Jiayuan Gu, Karl Pertsch, Oier Mees, Homer Rich Walke, Chuyuan Fu, Ishikaa Lunawat,IsabelSieh,SeanKirmani,etal. Evaluatingreal-worldrobotmanipulationpoliciesinsimulation. arXiv preprint arXiv:2405.05941, 2024. 2

  59. [61]

    Worldeval: World model as real-world robot policies evaluator.arXiv preprint arXiv:2505.19017, 2025

    Yaxuan Li, Yichen Zhu, Junjie Wen, Chaomin Shen, and Yi Xu. Worldeval: World model as real-world robot policies evaluator.arXiv preprint arXiv:2505.19017, 2025. 2

  60. [62]

    Genie envisioner: A unified world foundation platform for robotic manipulation.arXiv preprint arXiv:2508.05635, 2025

    Yue Liao, Pengfei Zhou, Siyuan Huang, Donglin Yang, Shengcong Chen, Yuxin Jiang, Yue Hu, Jingbin Cai, Si Liu, Jianlan Luo, et al. Genie envisioner: A unified world foundation platform for robotic manipulation.arXiv preprint arXiv:2508.05635, 2025. 19

  61. [63]

    Depth anything 3: Recovering the visual space from any views.arXiv preprint arXiv:2511.10647,

    Haotong Lin, Sili Chen, Junhao Liew, Donny Y Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang. Depth anything 3: Recovering the visual space from any views.arXiv preprint arXiv:2511.10647,

  62. [64]

    Truthfulqa: Measuring how models mimic human falsehoods

    Stephanie Lin, Jacob Hilton, and Owain Evans. Truthfulqa: Measuring how models mimic human falsehoods. InProceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers), pages 3214–3252, 2022. 2

  63. [65]

    Libero: Bench- marking knowledge transfer for lifelong robot learning.NeurIPS, 36:44776–44791, 2023

    Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. Libero: Bench- marking knowledge transfer for lifelong robot learning.NeurIPS, 36:44776–44791, 2023. 2, 4

  64. [66]

    Mamba4d: Efficient 4d point cloud video understanding with disentangled spatial-temporal state space models

    Jiuming Liu, Jinru Han, Lihao Liu, Angelica I Aviles-Rivero, Chaokang Jiang, Zhe Liu, and Hesheng Wang. Mamba4d: Efficient 4d point cloud video understanding with disentangled spatial-temporal state space models. InCVPR, pages 17626–17636, 2025. 3

  65. [67]

    Difflow3d: Hierarchical diffusion models for uncertainty-aware 3d scene flow estimation.TPAMI, 2025

    Jiuming Liu, Weicai Ye, Guangming Wang, Chaokang Jiang, Lei Pan, Jinru Han, Zhe Liu, Guofeng Zhang, and Hesheng Wang. Difflow3d: Hierarchical diffusion models for uncertainty-aware 3d scene flow estimation.TPAMI, 2025. 3

  66. [68]

    Arflow: Auto-regressive optical flow estimation for arbitrary-length videos via progressive next-frame forecasting

    Jiuming Liu, Mengmeng Liu, Siting Zhu, Yunpeng Zhang, Jiangtao Li, Michael Ying Yang, Francesco Nex, Hao Cheng, and Hesheng Wang. Arflow: Auto-regressive optical flow estimation for arbitrary-length videos via progressive next-frame forecasting. InICLR, 2026. 3

  67. [69]

    Towards interactive video world modeling: Frontiers, challenges, benchmarks, and future trends.arXiv preprint arXiv:2606.01164, 2026

    Jiuming Liu, Chaojun Ni, Mengmeng Liu, Chensheng Peng, Fangjinhua Wang, Sitian Shen, Marc Pollefeys, Masayoshi Tomizuka, Ayush Tewari, and Per Ola Kristensson. Towards interactive video world modeling: Frontiers, challenges, benchmarks, and future trends.arXiv preprint arXiv:2606.01164, 2026. 2

  68. [70]

    Robotransfer: Controllable geometry-consistent video diffusion for manipulation policy transfer

    Liu Liu, Xiaofeng Wang, Guosheng Zhao, Keyu Li, Wenkang Qin, Jiagang Zhu, Jiaxiong Qiu, Guan Huang, and Zhizhong Su. Robotransfer: Controllable geometry-consistent video diffusion for manipulation policy transfer. InCVPR, pages 1410–1420, 2026. 3 36 GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

  69. [71]

    Physgen: Rigid-body physics- grounded image-to-video generation

    Shaowei Liu, Zhongzheng Ren, Saurabh Gupta, and Shenlong Wang. Physgen: Rigid-body physics- grounded image-to-video generation. InECCV, pages 360–378. Springer, 2024. 2

  70. [72]

    Beyond fvd: Enhanced evaluation metrics for video generation quality, 2024

    Ge Ya Luo, Gian Mario Favero, Zhi Hao Luo, Alexia Jolicoeur-Martineau, and Christopher Pal. Beyond fvd: Enhanced evaluation metrics for video generation quality, 2024. URLhttps://arxiv.org/abs/ 2410.05203. 7

  71. [73]

    Viva: A video-generative value model for robot reinforcement learning.arXiv preprint arXiv:2604.08168, 2026

    Jindi Lv, Hao Li, Jie Li, Yifei Nie, Fankun Kong, Yang Wang, Xiaofeng Wang, Zheng Zhu, Chaojun Ni, Qiuping Deng, Hengtao Li, Jiancheng Lv, and Guan Huang. Viva: A video-generative value model for robot reinforcement learning.arXiv preprint arXiv:2604.08168, 2026. 3

  72. [74]

    Follow your pose: Pose-guided text-to-video generation using pose-free videos

    Yue Ma, Yingqing He, Xiaodong Cun, Xintao Wang, Siran Chen, Xiu Li, and Qifeng Chen. Follow your pose: Pose-guided text-to-video generation using pose-free videos. InAAAI, volume 38, pages 4117–4125, 2024a. 3

  73. [75]

    Follow-your-emoji: Fine-controllable and expressive freestyle portrait animation

    Yue Ma, Hongyu Liu, Hongfa Wang, Heng Pan, Yingqing He, Junkun Yuan, Ailing Zeng, Chengfei Cai, Heung-Yeung Shum, Wei Liu, et al. Follow-your-emoji: Fine-controllable and expressive freestyle portrait animation. InSIGGRAPH Asia, pages 1–12, 2024b

  74. [76]

    Controllable video generation: A survey.arXiv preprint arXiv:2507.16869, 2025a

    Yue Ma, Kunyu Feng, Zhongyuan Hu, Xinyu Wang, Yucheng Wang, Mingzhe Zheng, Xuanhua He, Chenyang Zhu, Hongyu Liu, Yingqing He, et al. Controllable video generation: A survey.arXiv preprint arXiv:2507.16869, 2025a

  75. [77]

    Follow-your-motion: Video motion transfer via efficient spatial-temporal decoupled finetuning.arXiv preprint arXiv:2506.05207, 2025b

    Yue Ma, Yulong Liu, Qiyuan Zhu, Ayden Yang, Kunyu Feng, Xinhua Zhang, Zhifeng Li, Sirui Han, Chenyang Qi, and Qifeng Chen. Follow-your-motion: Video motion transfer via efficient spatial-temporal decoupled finetuning.arXiv preprint arXiv:2506.05207, 2025b

  76. [78]

    Follow-your-emoji-faster: Towards efficient, fine-controllable, and expressive freestyle portrait animation.arXiv preprint arXiv:2509.16630, 2025c

    Yue Ma, Zexuan Yan, Hongyu Liu, Hongfa Wang, Heng Pan, Yingqing He, Junkun Yuan, Ailing Zeng, Chengfei Cai, Heung-Yeung Shum, et al. Follow-your-emoji-faster: Towards efficient, fine-controllable, and expressive freestyle portrait animation.arXiv preprint arXiv:2509.16630, 2025c

  77. [79]

    Follow-your-click: Open-domain regional image animation via motion prompts

    Yue Ma, Yingqing He, Hongfa Wang, Andong Wang, Leqi Shen, Chenyang Qi, Jixuan Ying, Chengfei Cai, Zhifeng Li, Heung-Yeung Shum, et al. Follow-your-click: Open-domain regional image animation via motion prompts. InAAAI, volume 39, pages 6018–6026, 2025d

  78. [80]

    Follow-your-creation: Empowering 4d creation through video inpainting.arXiv preprint arXiv:2506.04590, 2025e

    Yue Ma, Kunyu Feng, Xinhua Zhang, Hongyu Liu, David Junhao Zhang, Jinbo Xing, Yinhan Zhang, Ayden Yang, Zeyu Wang, and Qifeng Chen. Follow-your-creation: Empowering 4d creation through video inpainting.arXiv preprint arXiv:2506.04590, 2025e

  79. [81]

    Group editing: Edit multiple images in one go.arXiv preprint arXiv:2603.22883, 2026a

    Yue Ma, Xinyu Wang, Qianli Ma, Qinghe Wang, Mingzhe Zheng, Xiangpeng Yang, Hao Li, Chongbo Zhao, Jixuan Ying, Harry Yang, et al. Group editing: Edit multiple images in one go.arXiv preprint arXiv:2603.22883, 2026a

  80. [82]

    Fastvmt: Eliminating redundancy in video motion transfer.arXiv preprint arXiv:2602.05551, 2026b

    Yue Ma, Zhikai Wang, Tianhao Ren, Mingzhe Zheng, Hongyu Liu, Jiayi Guo, Mark Fong, Yuxuan Xue, Zixiang Zhao, Konrad Schindler, et al. Fastvmt: Eliminating redundancy in video motion transfer.arXiv preprint arXiv:2602.05551, 2026b. 3

Showing first 80 references.