FuguReport

Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation

Authors Atin Pothiraj, Jaemin Cho, Yue Zhang, Elias Stengel-Eskin, Mohit Bansal
Affiliations The University of North Carolina at Chapel Hill / The University of Texas at Austin / AI2 / The Johns Hopkins University
Categories Evaluation / Physical Plausibility Evaluation / Hierarchical question-based video assessment, Method / Scene Graph Analysis / Object, action, and physics consistency checks, Application / Text-to-Video Generation / Evaluating physics fidelity in generated videos
License CC BY 4.0

Abstract Overview

This paper introduces Physics Question Scene Graph (PQSG), a hierarchical evaluation framework for assessing whether text-to-video outputs are faithful to prompts at the levels of objects, actions, and physical plausibility. PQSG represents evaluation as a directed acyclic graph of atomic verification questions with explicit dependencies, so physics questions are asked only when prerequisite object and action conditions are satisfied. To validate the framework, the authors build FinePhyEval, a human-annotated benchmark of physics-focused prompts and generated videos from recent video generation models. Experiments show that PQSG provides both aggregate scoring and fine-grained failure localization, with stronger correlation to human judgments than several prior evaluation baselines.

Novelty

The paper's main novelty is a graph-structured, question-based evaluation pipeline that explicitly models logical dependencies between object, action, and physics checks for generated videos. It also contributes FinePhyEval, a benchmark with human annotations for both overall video quality categories and the underlying question-generation and question-answering subtasks.

Results

On FinePhyEval, PQSG achieves higher correlation with human overall judgments than the compared baselines, reaching Pearson correlations up to 0.478 with GPT-5.5-based QA and 0.80 when human QA is used. The framework ranks Sora 2 and Veo 3 above Wan 2.1 and Cosmos 2.5 on physical realism, and shows that current VLMs are strong at question generation but notably weaker at answering physics-focused questions, with the best reported physics QA accuracy at 64.6%.

Key Points

  1. PQSG decomposes video evaluation into object, action, and physics questions linked by dependency constraints, which helps localize failure modes and avoid invalid downstream judgments.
  2. FinePhyEval provides 195 human-rated videos for Likert-scale evaluation and additional fine-grained annotations for question generation and question answering analysis.
  3. The experiments indicate that object depiction is easier for current generators and evaluators than action and physics, and that the dependency graph and fine-grained questioning both improve alignment with human ratings.

References

This page was created using generative AI such as GPT-5, Claude Opus 4, Gemini 3, Gemini 3.1 Flash Image, and their higher-end successor versions. No guarantee can be made regarding its contents.