論文の概要: PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop
- arxiv url: http://arxiv.org/abs/2610.00559v1
- Date: Wed, 30 Sep 2026 18:34:22 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-03 01:19:23.707444
- Title: PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop
- Title(参考訳): PhysVista:知覚推論評価ループによるVLMの物理知能のベンチマーク
- Abstract要約: 視覚言語モデル(VLM)における物理インテリジェンスを評価するためのベンチマークであるPhysVistaを紹介する。
PhysVistaは、人間の観察・推論プロセスにインスパイアされた、クローズドな認知ループの枠組みを復元する。
さらにPhysVistaには、現実世界とAI生成ビデオの両方が組み込まれている。
- 参考スコア(独自算出の注目度): 23.759532799031334
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Vision-Language Models (VLMs) have shown strong multimodal reasoning capabilities, yet whether they truly capture the physical consistency underlying real-world dynamics remains unclear. Existing benchmark paradigms often suffer from fragmented evaluation, focusing on isolated cognitive stages while overlooking the inherent synergy between perception, reasoning, and physical judgment. The lack of a holistic perspective limits the ability to diagnose whether VLMs can reliably evaluate the physical authenticity of emerging generative models. To address these issues, we introduce PhysVista, a benchmark designed to evaluate physical intelligence in VLMs through a closed cognitive loop framework inspired by the human seeing-reasoning-assessment process. PhysVista restores this loop by jointly evaluating physical state perception, physical dynamics reasoning, and physical plausibility assessment. It further distinguishes event-level reasoning and scale-level reasoning to enable fine-grained analysis of physical understanding. In addition, PhysVista incorporates both real-world and AI-generated videos, allowing evaluation across diverse domains and emerging generative scenarios. Extensive experiments across a diverse set of VLMs reveal substantial limitations in physical reasoning and plausibility assessment, highlighting a persistent gap between visual recognition and genuine physical understanding, and pointing toward more principled designs for physically grounded multimodal intelligence.
- Abstract(参考訳): VLM(Vision-Language Models)は、強力なマルチモーダル推論能力を示しているが、実世界の力学の基盤となる物理的な一貫性を真に捉えているかどうかは不明だ。
既存のベンチマークパラダイムは、しばしば断片的な評価に悩まされ、独立した認知段階に焦点を合わせながら、知覚、推論、物理的判断の固有の相乗効果を見越す。
全体論的視点の欠如は、VLMが新しい生成モデルの物理的信頼性を確実に評価できるかどうかを診断する能力を制限する。
これらの問題に対処するため,人間の視線推論プロセスにインスパイアされたクローズド認知ループフレームワークを用いて,VLMの物理的インテリジェンスを評価するためのベンチマークであるPhysVistaを紹介した。
PhysVistaはこのループを、物理的状態認識、物理力学推論、物理的妥当性評価と共同で評価することで復元する。
さらに、イベントレベルの推論とスケールレベルの推論を区別し、物理的理解のきめ細かい分析を可能にする。
さらにPhysVistaには、現実世界とAI生成ビデオの両方が組み込まれている。
多様なVLMの集合にわたる広範囲な実験は、物理的推論と可視性評価の重大な制限を明らかにし、視覚認識と真の物理的理解の間に永続的なギャップを浮き彫りにし、物理的に基礎付けられたマルチモーダルインテリジェンスのためのより原則化された設計を指している。
関連論文リスト
- PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning [10.8081248403127]
ビデオ言語モデル(VLM)は,映像理解と視覚的質問応答において優れた性能を発揮している。
トレーニング不要な物理メモリと検証フレームワークであるPhysMRVを提案する。
論文 参考訳(メタデータ) (2026-07-11T08:06:04Z) - NICE FACT: Diagnosing and Calibrating VLMs in Quantitative Reasoning for Kinematic Physics [65.02899948986969]
この研究は、視覚言語モデルが物理的世界をどのように知覚するかを根本的に理解し、物理法則を活用することを目的としている。
運動物理学の量的推論を明示的に分解する双対診断パラダイムであるNICEとFACTを提案する。
NICEは、我々の新しい地区インフォームドキャリブレーション手法と、信頼性の評価とキャリブレーションのための新しいメトリクスについて研究する。
論文 参考訳(メタデータ) (2026-05-08T20:17:44Z) - EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving [40.67852883277799]
VLM(Vision-Language Models)は、自律運転における高度な推論である。
しかし、この推論を根底にあるエゴ運動の物理学に根ざす能力は、いまだに理解されていない。
視覚中心基礎モデルのセマンティック・エゴモーション理解を評価するための診断ベンチマークであるEgoDyn-Benchを紹介する。
論文 参考訳(メタデータ) (2026-04-22T07:49:02Z) - Does Physics Knowledge Emerge in Frontier Models? [19.035965618393096]
VLM(Leading Vision-Language Models)は、視覚知覚と一般的な推論において強力な結果を示す。
しかし、物理力学を理解し予測する能力は、まだ不明である。
3つの物理シミュレーションデータセット上で6つのフロンティアVLMをベンチマークする。
論文 参考訳(メタデータ) (2025-10-03T22:30:06Z) - ContPhy: Continuum Physical Concept Learning and Reasoning from Videos [86.63174804149216]
ContPhyは、マシン物理常識を評価するための新しいベンチマークである。
私たちは、さまざまなAIモデルを評価し、ContPhyで満足なパフォーマンスを達成するのに依然として苦労していることがわかった。
また、近年の大規模言語モデルとパーティクルベースの物理力学モデルを組み合わせるためのオラクルモデル(ContPRO)を導入する。
論文 参考訳(メタデータ) (2024-02-09T01:09:21Z) - Intrinsic Physical Concepts Discovery with Object-Centric Predictive
Models [86.25460882547581]
PHYsical Concepts Inference NEtwork (PHYCINE) は、異なる抽象レベルの物理概念を監督なしで推論するシステムである。
物理概念変数を含むオブジェクト表現は因果推論タスクの性能向上に有効であることを示す。
論文 参考訳(メタデータ) (2023-03-03T11:52:21Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。