Fugu-MT 論文翻訳(概要): WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors

論文の概要: WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors

arxiv url: http://arxiv.org/abs/2605.10434v1
Date: Mon, 11 May 2026 12:06:57 GMT
ステータス: 翻訳完了
システム内更新日: 2026-05-12 23:28:50.791688
Title: WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors
Title（参考訳）: WorldReasonBench:将来的な世界予測者としてのビデオジェネレータの人為的なストレステスト
Authors: Keming Wu, Yijing Cui, Wenhan Xue, Qijie Wang, Xuan Luo, Zhiyuan Feng, Zuhao Yang, Sudong Wang, Sicong Jiang, Haowei Zhu, Zihan Wang, Ping Nie, Wenhu Chen, Bin Wang,
Abstract要約: 本稿では,映像生成評価を世界状態予測として再設定するWorldReasonBenchを紹介する。人手による2部構成手法を用いて生成した映像の評価を行った。 WorldRewardBenchは、約6Kのエキスパートアノテートされたペアが1.4Kビデオに対して設定された選好ベンチマークである。
参考スコア（独自算出の注目度）: 45.545823511469166
License: http://creativecommons.org/licenses/by/4.0/
Abstract: Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving into "world simulators." Yet the community still lacks a benchmark that directly tests whether a model can reason about how an observed world should evolve over time. We introduce WorldReasonBench, which reframes video generation evaluation as world-state prediction: given an initial state and an action, can a model generate a future video whose state evolution remains physically, socially, logically, and informationally consistent? WorldReasonBench contains 436 curated test cases with structured ground-truth QA annotations spanning four reasoning dimensions and 22 subcategories. We evaluate generated videos with a human-aligned two-part methodology: Process-aware Reasoning Verification uses structured QA and reasoning-phase diagnostics to detect temporal and causal failures, while Multi-dimensional Quality Assessment scores reasoning quality, temporal consistency, and visual aesthetics for ranking and reward modeling. We further introduce WorldRewardBench, a preference benchmark with approximately 6K expert-annotated pairs over 1.4K videos, supporting pair-wise and point-wise reward-model evaluation. Across modern video generators, our results expose a persistent gap between visual plausibility and world reasoning: videos can look convincing while failing dynamics, causality, or information preservation. We will release our benchmarks and evaluation toolkit to support community research on genuinely world-aware video generation at https://github.com/UniX-AI-Lab/WorldReasonBench/.
Abstract（参考訳）: Seedance2.0やVeo3.1のような商用ビデオ生成システムは急速に改善され、ビデオジェネレータが「世界シミュレータ」に進化しつつあるという見方が強まった。しかし、コミュニティはまだ、観察された世界が時間とともにどのように進化すべきかをモデルが推論できるかどうかを直接テストするベンチマークを欠いている。初期状態とアクションが与えられたら、モデルは、物理的、社会的、論理的、情報的に整合した状態の将来のビデオを生成することができるか? WorldReasonBenchには、4つの推論次元と22のサブカテゴリにまたがる構造化された接地型QAアノテーションを備えた436のキュレートされたテストケースが含まれている。プロセス認識推論検証では、構造化されたQAと推論フェーズの診断を用いて時間的・因果的故障を検知し、多次元品質評価では、品質、時間的一貫性、視覚的美学を評価・評価する。さらに、約6Kのエキスパートアノテートペアを1.4Kビデオ上に配置し、ペアワイドとポイントワイドの報酬モデル評価をサポートする選好ベンチマークであるWorldRewardBenchを紹介する。現代のビデオジェネレータでは、ビデオはダイナミックス、因果性、情報保存に失敗しながら説得力のあるように見える。我々はベンチマークと評価ツールキットをリリースし、https://github.com/UniX-AI-Lab/WorldReasonBench/で真に世界対応のビデオ生成に関するコミュニティリサーチをサポートする。

論文の概要: WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors

関連論文リスト