PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
Abstract Overview
This paper argues that evaluating video generators as visual world models requires assessing the distribution of possible futures under a fixed observation and action, rather than judging only single-rollout plausibility. It formalizes this criterion as probabilistic alignment and introduces PAWBench, a 50-scenario benchmark across eight physical mechanism groups, alongside PAWEval, a rubric-based protocol mapping repeated rollouts to empirical outcome distributions. PAWBench separates probability-mass alignment (PAW-Calibration) from valid-support recovery (PAW-Coverage) to support evaluation when reference probabilities are known versus when only valid outcomes can be enumerated. Evaluating eleven video generation models reveals that current systems fall substantially short of probabilistically aligned world modeling.
Novelty
The paper defines probabilistic alignment as a distribution-level requirement for video world models and operationalizes it via repeated rollouts under identical initial observations and actions. It also establishes the two-track PAWBench and PAWEval framework, which decouples testing calibration against known reference probabilities from evaluating coverage over valid outcome supports.
Results
Across 50 scenarios and 11 video generators, no model simultaneously achieves accurate outcome probabilities, broad valid-support coverage, and high scene-level reliability. Controlled tests show that models underreact to causal physical alterations while being swayed by non-causal text distractors. Furthermore, while prompt engineering, coupled noise sampling, and training data adjustments can steer outputs or broaden exploration, none reliably recovers correct scene-conditioned future distributions.
Key Points
- PAWBench evaluates repeated video rollouts under fixed initial conditions, decoupling probability calibration from valid future coverage across 50 physical scenarios.
- Empirical benchmarks across eleven video generation models demonstrate that single-video plausibility does not guarantee aligned future distributions.
- Interventions spanning language prompting, initial noise coupling, and parameter fine-tuning fail to reliably produce calibrated, scene-dependent future distributions.
References
- arXiv: https://arxiv.org/abs/2608.27345v1
- Fugu-MT: https://fugumt.com/fugumt/paper_check/2608.27345v1
- Hugging Face Papers: https://huggingface.co/papers/2608.27345
- Project: https://pawbench.github.io