REVIEW 2 major objections 60 references
A new benchmark reveals that current interactive world models lack reliable long-horizon stability across multiple dimensions.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-07-03 21:56 UTC pith:KWMJVSH3
load-bearing objection The paper defines four new metric families for long-horizon world model stability but supplies no correlation checks or representativeness stats to back them up. the 2 major comments →
WorldOdysseyBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Existing benchmarks evaluate action following only at the trajectory level while ignoring memory and interaction physics. WorldOdysseyBench addresses this with four dimensions: per-frame action metric to expose hidden failures, segment-based drift for non-monotonic collapse, controllability-gated physics for plausibility, and action-decoupled memory via 3D reconstruction and tracking. Evaluation of over ten models finds none reliably satisfies all, with the best scoring only moderately on the 600+ test cases in nature, urban, and indoor environments.
What carries the argument
WorldOdysseyBench benchmark with its four specialized metrics for per-frame action accuracy, segment-based visual drift, controllability-gated physics plausibility, and action-decoupled memory assessment through 3D point clouds and tracking.
Load-bearing premise
The four metrics correctly measure the intended aspects of long-horizon stability and the test cases represent open-world interactions.
What would settle it
A model that scores high across all four metrics yet fails in extended real interactive tasks, or a model that fails the benchmark metrics but succeeds in sustained real-world use.
If this is right
- Models that improve on the four metrics would support more reliable deployment in applications requiring continuous interaction.
- Trajectory-level scores can conceal per-frame action mistakes and sudden visual collapses in the middle of sequences.
- Physics checks must be gated on accurate action execution to isolate true physical consistency from control errors.
- Memory tests performed without action cues can separately verify recall of scenes and tracked subjects.
Where Pith is reading between the lines
- The benchmark design could push development toward models that combine learned prediction with separate memory buffers.
- Longer test sequences beyond 60 seconds might expose additional memory decay patterns not visible in the current cases.
- Direct transfer of high-scoring models to physical robot control could test whether benchmark results predict real stability.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces WorldOdysseyBench, an open-world benchmark for long-horizon stability of interactive world models (IWMs). It defines four tailored metrics—per-frame action (bypassing semantic scale issues), segment-based drift for vision (capturing mid-sequence collapse), controllability-gated physics (over mechanics/optics/3D), and action-decoupled memory (via 3D reconstruction and tracking+VLM)—evaluated on 600+ test cases spanning Nature/Urban/Indoor scenes in first/third-person views with 10-60s WASD interactions. Evaluation of 10+ open/closed-source models finds none reliably satisfy all dimensions, with the best achieving only moderate scores.
Significance. If the metrics prove valid and the test cases representative, the benchmark would be a useful contribution by exposing stability, physics, and memory failures missed by prior trajectory-level or start-vs-end evaluations, providing a concrete framework to drive progress toward deployable IWMs.
major comments (2)
- [Abstract] Abstract: the central claim that the four metrics 'capture failures missed by prior trajectory-level or start-vs-end metrics' and that 'none reliably satisfies all dimensions' is load-bearing on the unverified assumption that per-frame action, segment-based drift, controllability-gated physics, and action-decoupled memory correctly quantify the intended aspects; no correlation study, ablation against existing metrics, or human-judgment validation is referenced.
- [Abstract] Abstract: the representativeness of the 600+ test cases for open-world interaction is asserted without supporting diversity statistics (e.g., scene-type distribution, duration histogram, viewpoint balance), which directly affects whether the headline result generalizes beyond the chosen cases.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback. We address each major comment below and will revise the manuscript accordingly to strengthen the presentation of the benchmark.
read point-by-point responses
-
Referee: [Abstract] Abstract: the central claim that the four metrics 'capture failures missed by prior trajectory-level or start-vs-end metrics' and that 'none reliably satisfies all dimensions' is load-bearing on the unverified assumption that per-frame action, segment-based drift, controllability-gated physics, and action-decoupled memory correctly quantify the intended aspects; no correlation study, ablation against existing metrics, or human-judgment validation is referenced.
Authors: The metrics were designed to target concrete, observable limitations of prior evaluations, as motivated in the introduction: per-frame action avoids semantic scale disparities that affect trajectory-level scores, segment-based drift detects mid-sequence non-monotonic collapses invisible to start-vs-end comparisons, controllability-gated physics isolates physical plausibility under correct action execution, and action-decoupled memory separates scene and subject recall from action following. The empirical results across 10+ models illustrate distinct failure patterns across these axes that standard metrics do not surface. We agree, however, that explicit supporting analyses were not included in the submission. In the revision we will add an appendix containing (i) an ablation comparing the new metrics against trajectory-level action accuracy, FVD, and start-vs-end drift on the same model outputs and (ii) a small-scale human preference study on a 50-case subset to check alignment with perceived stability, physics, and memory errors. revision: yes
-
Referee: [Abstract] Abstract: the representativeness of the 600+ test cases for open-world interaction is asserted without supporting diversity statistics (e.g., scene-type distribution, duration histogram, viewpoint balance), which directly affects whether the headline result generalizes beyond the chosen cases.
Authors: We agree that quantitative diversity statistics would make the claim of representativeness more transparent. The test cases were collected to span Nature/Urban/Indoor environments, first- and third-person viewpoints, and 10–60 s durations, but these were described qualitatively rather than with explicit distributions. In the revised manuscript we will add a dedicated subsection (or table) in the benchmark description reporting scene-type counts, a duration histogram, and first/third-person balance across the full set of 600+ cases. revision: yes
Circularity Check
No significant circularity; benchmark definition paper
full rationale
This is a benchmark construction paper that defines four new evaluation metrics (per-frame action, segment-based drift, controllability-gated physics, action-decoupled memory) and applies them to existing models. No equations, parameter fits, or predictions are present that could reduce to the inputs by construction. The central claims rest on the explicit definitions of the metrics and the collection of 600+ test cases rather than any self-referential derivation or load-bearing self-citation chain. The evaluation results follow directly from applying the stated protocols and do not presuppose the outcomes.
Axiom & Free-Parameter Ledger
read the original abstract
Despite rapid progress in interactive world models (IWMs), existing benchmarks evaluate action following only at trajectory level and ignore memory and interaction physics. We introduce WorldOdysseyBench, an open-world benchmark for long-horizon stability across four dimensions, each with tailored innovations: (i) Action: per-frame action metric bypassing cross-model semantic scale disparity and exposing failures hidden by trajectory; (ii) Vision: segment-based drift metric capturing non-monotonic mid-sequence collapse missed by start-vs-end comparisons; (iii) Physics: controllability-gated evaluation over mechanics, optics, and 3D consistency, scoring plausibility under faithful action execution; (iv) Memory: action-decoupled protocol evaluating scene memory via transition-localized 3D point-cloud reconstruction and subject memory via tracking-plus-VLM reasoning. The benchmark comprises 600+ test cases across Nature, Urban, and Indoor scenes in first/third-person views with WASD 10-60s continuous interaction. Evaluating 10+ open/closed-source models reveals none reliably satisfies all dimensions; even the best achieves only moderate scores. Advances on WorldOdysseyBench are steps toward IWMs that are stable, physically grounded, memory-faithful, and deployable in real-world applications.
Figures
Reference graph
Works this paper leans on
-
[1]
Genie: Generative interactive environments
Bruce, J., Dennis, M., Edwards, A., et al. Genie: Generative interactive environments. InICML, 2024
work page 2024
-
[2]
Google DeepMind. Genie 3: A new frontier for world models.https://deepmind.google/discover/ blog/genie- 3- a- new- frontier- for- world- models/, 2025
work page 2025
-
[3]
Alibaba Group. Happy Oyster: An open-ended world model for real-time world creation and interaction.https:// happyoyster.cn/, 2026
work page 2026
-
[4]
Wan: Open and Advanced Large-Scale Video Generative Models
Wan Team. Wan 2.1: A comprehensive and unified video generation model.arXiv preprint arXiv:2503.20314, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[5]
Kling 3.0: Next-generation AI video generation
Kuaishou. Kling 3.0: Next-generation AI video generation. https://klingai.com/, 2025
work page 2025
-
[6]
Sora 2: A large-scale video generation model
OpenAI. Sora 2: A large-scale video generation model. https://openai.com/sora/, 2025
work page 2025
-
[7]
Veo 3: State-of-the-art video generation
Google DeepMind. Veo 3: State-of-the-art video generation. https : / / deepmind . google / technologies / veo/, 2025
work page 2025
-
[8]
Monkeyocr: Document parsing with a structure-recognition-relation triplet paradigm, 2026
ByteDance Seed Team. Seedance 2.0: Advancing video generation for world complexity.arXiv preprint arXiv:2506.05218, 2025
-
[9]
Matrix-game 2.0: An open-source real-time and streaming interactive world model
Matrix-Game Team. Matrix-Game 2.0: An open-source, real-time, and streaming interactive world model.arXiv preprint arXiv:2508.13009, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[10]
Matrix-Game Team. Matrix-Game 3.0: Real-time and streaming interactive world model with long-horizon mem- ory.arXiv preprint, 2026
work page 2026
-
[11]
HY-World Team. HY-World 1.5: A systematic framework for interactive world modeling with real-time latency and ge- ometric consistency.arXiv preprint, 2025
work page 2025
-
[12]
Yume 1.5: A text-controlled interactive world generation model.arXiv preprint, 2026
Yume Team. Yume 1.5: A text-controlled interactive world generation model.arXiv preprint, 2026
work page 2026
-
[13]
Advancing open-source world models.arXiv preprint, 2026
LingBot Team. Advancing open-source world models.arXiv preprint, 2026. 16
work page 2026
-
[14]
arXiv preprint arXiv:2602.08025 (2026)
Ye, H., Lu, J., et al. MIND: Benchmarking memory consis- tency and action following in world models.arXiv preprint arXiv:2602.08025, 2026
-
[15]
WorldMark: A Unified Benchmark Suite for Interactive Video World Models
Alaya Studio. WorldMark: A unified benchmark suite for interactive video world models.arXiv preprint arXiv:2604.21686, 2026
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[16]
iWorld-Bench: A benchmark for interactive world models with a unified action generation framework
Li, Y ., et al. iWorld-Bench: A benchmark for interactive world models with a unified action generation framework. InICML, 2026
work page 2026
-
[17]
Shanda AI. WildWorld: A large-scale dataset for dynamic world modeling with actions and explicit state toward gener- ative ARPG.arXiv preprint arXiv:2603.23497, 2026
-
[18]
VBench: Comprehensive benchmark suite for video generative mod- els
Huang, Z., He, Y ., Yu, J., Zhang, F., Si, C., et al. VBench: Comprehensive benchmark suite for video generative mod- els. InCVPR, 2024
work page 2024
-
[19]
Huang, Z., Zhang, F., Xu, X., He, Y ., Yu, J., et al. VBench++: Comprehensive and versatile benchmark suite for video gen- erative models.arXiv preprint arXiv:2411.13503, 2024
-
[20]
VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
Zheng, D., Huang, Z., Liu, H., Zou, K., He, Y ., et al. VBench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness.arXiv preprint arXiv:2503.21755, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[21]
WorldScore: A unified evaluation benchmark for world generation
Stanford. WorldScore: A unified evaluation benchmark for world generation. InICCV, 2025
work page 2025
-
[22]
WorldModelBench: Judging video generation models as world models
UC Berkeley. WorldModelBench: Judging video generation models as world models. InNeurIPS, 2025
work page 2025
-
[23]
Towards world simulator: Crafting physical commonsense-based benchmark for video generation
SJTU, et al. Towards world simulator: Crafting physical commonsense-based benchmark for video generation. In ICML, 2025
work page 2025
-
[24]
UCLA. WorldBench: Disambiguating physics for di- agnostic evaluation of world models.arXiv preprint arXiv:2601.21282, 2026
-
[25]
Video PreTraining (VPT): Learning to act by watching unlabeled online videos
Baker, B., et al. Video PreTraining (VPT): Learning to act by watching unlabeled online videos. InNeurIPS, 2022
work page 2022
-
[26]
Microsoft. MineWorld: A real-time and open-source interactive world model on Minecraft.arXiv preprint arXiv:2504.08388, 2025
-
[27]
NVIDIA. SANA-WM: Efficient minute-scale world model- ing with hybrid linear diffusion transformer.arXiv preprint, 2026
work page 2026
-
[28]
Lyra 2.0: Explorable generative 3D worlds.arXiv preprint, 2026
NVIDIA. Lyra 2.0: Explorable generative 3D worlds.arXiv preprint, 2026
work page 2026
-
[29]
minWM Team. minWM: A full-stack open-source frame- work for real-time interactive video world models.arXiv preprint, 2026
work page 2026
-
[30]
LAION- 5B: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., et al. LAION- 5B: An open large-scale dataset for training next generation image-text models. InNeurIPS, 2022
work page 2022
-
[31]
MUSIQ: Multi-scale image quality transformer
Ke, J., Wang, Q., Wang, Y ., Milanfar, P., and Yang, F. MUSIQ: Multi-scale image quality transformer. InICCV, 2021
work page 2021
-
[32]
Helios: A comprehensive benchmark for video generative models.arXiv preprint, 2025
Helios Team. Helios: A comprehensive benchmark for video generative models.arXiv preprint, 2025
work page 2025
-
[33]
WorldCompass: Reinforcement learning for long-horizon world models.arXiv preprint, 2026
WorldCompass Team. WorldCompass: Reinforcement learning for long-horizon world models.arXiv preprint, 2026
work page 2026
-
[34]
ViPE: Visual pose estimation for camera trajectory recovery
-
[35]
Teed, Z. and Deng, J. DROID-SLAM: Deep visual SLAM for monocular, stereo, and RGB-D cameras. InNeurIPS, 2021
work page 2021
-
[36]
Ying, Z., et al. WBench: A comprehensive benchmark for evaluating world models via action-conditioned video gener- ation.arXiv preprint, 2026
work page 2026
-
[37]
Chen, Y . and Medioni, G. Object modelling by registra- tion of multiple range images.Image and Vision Computing, 10(3):145–155, 1992
work page 1992
-
[38]
Rusu, R. B., Blodow, N., and Beetz, M. Fast Point Feature Histograms (FPFH) for 3D registration. InICRA, 2009
work page 2009
-
[39]
Point Transformer V3: Simpler, faster, stronger
Wu, X., Jiang, L., Wang, P.-S., et al. Point Transformer V3: Simpler, faster, stronger. InCVPR, 2024
work page 2024
-
[40]
SAM 2: Segment Anything in Images and Videos
Ravi, N., Gabeur, V ., Hu, Y .-T., et al. SAM 2: Seg- ment anything in images and videos.arXiv preprint arXiv:2408.00714, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[41]
Two-frame motion estimation based on poly- nomial expansion
Farneb ¨ack, G. Two-frame motion estimation based on poly- nomial expansion. InScandinavian Conference on Image Analysis (SCIA), 2003
work page 2003
-
[42]
Bu, W., Wu, Y ., Yu, Q., et al. What limits virtual agent appli- cation? OmniBench: A scalable multi-dimensional bench- mark for essential virtual agent capabilities. InICML, 2025
work page 2025
-
[43]
Wu, M., Cai, Z., Zhao, F., et al. Omni-WorldBench: Towards a comprehensive interaction-centric evaluation for world models.arXiv preprint arXiv:2603.22212, 2026
-
[44]
Do vision-language models have internal world models? Towards an atomic evaluation
Gao, Q., Pi, X., Liu, K., et al. Do vision-language models have internal world models? Towards an atomic evaluation. InACL, 2025
work page 2025
-
[45]
How far is video generation from world model: A physical law perspective
Kang, B., Yue, Y ., Lu, R., et al. How far is video generation from world model: A physical law perspective. InICML, 2025
work page 2025
-
[46]
WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World
Liang, A., Kong, L., Yan, T., et al. WorldLens: Full-spectrum evaluations of driving world models in real world.arXiv preprint arXiv:2512.10958, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[47]
Shang, Y ., Li, Z., Ma, Y ., et al. WorldArena: A unified benchmark for evaluating perception and functional utility of embodied world models.arXiv preprint arXiv:2602.08971, 2026
-
[48]
Arai, H., Ishihara, K., Takahashi, T., and Yamaguchi, Y . ACT-Bench: Towards action controllable world models for autonomous driving.arXiv preprint arXiv:2412.05337, 2024
-
[49]
EWMBench: Evaluat- ing scene, motion, and semantic quality in embodied world models
Hu, Y ., Huang, S., Liao, Y ., et al. EWMBench: Evaluat- ing scene, motion, and semantic quality in embodied world models. InBMVC, 2025
work page 2025
-
[50]
WorldOlympiad: Can Your World Model Survive a Triathlon?
Zhao, Y ., Zhao, W., Wang, W., Zhang, Z., An, D., Liu, A., Yu, Y ., Tang, J., Wang, F., Wang, W., and Zhuang, B. Worl- dOlympiad: Can Your World Model Survive a Triathlon? arXiv preprint arXiv:2606.11129, 2026
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[51]
Person moves forward; Camera turns left
Teed, Z. and Deng, J. RAFT: Recurrent All-Pairs Field Transforms for Optical Flow.arXiv preprint arXiv:2003.12039, 2020. 17 A. Test Suite Gallery To provide a qualitative overview of the visual and action coverage in WorldOdysseyBench, we include a gallery of representative test cases in Figure 11. Each panel shows the first-frame image together with an a...
-
[52]
Depth-percentile Filteringretain nearest p%Far Near Cross-Segment Registration Coarse:PTY3 / FPFH+ RANSACFine:point-to-planeICPQuality-awareacceptance:fitness ≥ τᵩand Chamferimprovest … Video + GTActions⋯WW(Explore) SS(Revisit)⋯ Observation Segment(t ≤ t*) Frame-LevelPointClouds t Revisit Segment(t > t*) t Frame-LevelPointClouds t*Memory Evaluation Pipeli...
-
[53]
Holistic scoring accommodates smooth viewpoint and illumination changes that can confound frame-level com- parisons, while producing a single benchmark-compatible scalar without requiring an external reference-feature li- brary. G. Subject Memory Prompt For third-person memory evaluation, the Subject Memory Evaluation Prompt in Figure 27 scores video-leve...
-
[54]
Is there a shadow visible on the wall or floor?
-
[55]
If yes, does the shadow move or change in a way that is physically consistent with the camera movement and the light source position? Answer ’yes’ if the shadow appears and behaves correctly, ’no’ if the shadow is missing or behaves incorrectly. After your reasoning, conclude with exactly one line: Answer: yes or Answer: no Figure 26.Occlusion and shadow ...
-
[56]
Identity change: the subject becomes a different individual or category
-
[57]
Structural distortion: the body, anatomy, proportions, limbs, head, face, or key parts become deformed or implausible
-
[58]
Appearance drift: color, texture, clothing, hair, material, or style is truly rewritten
-
[59]
Subject disappearance: the subject becomes partly or fully invisible
-
[60]
Quality degradation: severe blur, low resolution, diffused edges, or loss of key details makes the subject hard to identify. Important calibration rules: - Viewpoint change alone is not inconsistency. Back-to-front, front-to-back, side-to-front, or far-to-close changes are acceptable if the subject can reasonably be the same individual under the new view....
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.