FuguReport

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

Authors Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram Đorđević, Shiyang Li, Yifan Zhou, Bin Fu, Wenlong Zhang, Junjun He, Yu Qiao, Yihao Liu, Jingbo Xing, Xi Chen
Affiliations Shanghai Jiao Tong University / Huggingface / The University of Hong Kong / Shanghai Innovation Institute / Shanghai AI Laboratory / Krea AI / Alibaba
Categories Evaluation / Benchmarking / Probabilistic alignment evaluation, Method / World Modeling / Video rollout empirical distribution, Application / Video Generation / Evaluating world dynamics samples
License CC BY 4.0

Abstract Overview

This paper argues that evaluating video generators as visual world models requires assessing the distribution of possible futures under a fixed observation and action, rather than judging only single-rollout plausibility. It formalizes this criterion as probabilistic alignment and introduces PAWBench, a 50-scenario benchmark across eight physical mechanism groups, alongside PAWEval, a rubric-based protocol mapping repeated rollouts to empirical outcome distributions. PAWBench separates probability-mass alignment (PAW-Calibration) from valid-support recovery (PAW-Coverage) to support evaluation when reference probabilities are known versus when only valid outcomes can be enumerated. Evaluating eleven video generation models reveals that current systems fall substantially short of probabilistically aligned world modeling.

Novelty

The paper defines probabilistic alignment as a distribution-level requirement for video world models and operationalizes it via repeated rollouts under identical initial observations and actions. It also establishes the two-track PAWBench and PAWEval framework, which decouples testing calibration against known reference probabilities from evaluating coverage over valid outcome supports.

Results

Across 50 scenarios and 11 video generators, no model simultaneously achieves accurate outcome probabilities, broad valid-support coverage, and high scene-level reliability. Controlled tests show that models underreact to causal physical alterations while being swayed by non-causal text distractors. Furthermore, while prompt engineering, coupled noise sampling, and training data adjustments can steer outputs or broaden exploration, none reliably recovers correct scene-conditioned future distributions.

Key Points

  1. PAWBench evaluates repeated video rollouts under fixed initial conditions, decoupling probability calibration from valid future coverage across 50 physical scenarios.
  2. Empirical benchmarks across eleven video generation models demonstrate that single-video plausibility does not guarantee aligned future distributions.
  3. Interventions spanning language prompting, initial noise coupling, and parameter fine-tuning fail to reliably produce calibrated, scene-dependent future distributions.

References

This page was created using generative AI such as GPT-5, Claude Opus 4, Gemini 3, Gemini 3.1 Flash Image, and their higher-end successor versions. No guarantee can be made regarding its contents.