論文の概要: PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning
- arxiv url: http://arxiv.org/abs/2607.10190v1
- Date: Sat, 11 Jul 2026 08:06:04 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-14 15:40:48.333323
- Title: PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning
- Title(参考訳): PhysMRV:物理記憶の検索と物理可視性推論の検証
- Authors: Wenyuan Wang, Lianyu Hu, Hao Wang, Yang Liu,
- Abstract要約: ビデオ言語モデル(VLM)は,映像理解と視覚的質問応答において優れた性能を発揮している。
トレーニング不要な物理メモリと検証フレームワークであるPhysMRVを提案する。
- 参考スコア(独自算出の注目度): 10.8081248403127
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Video-language models (VLMs) have achieved remarkable performance on video understanding and visual question answering, yet they remain unreliable in reasoning about physical plausibility, where understanding object interactions, causal dynamics, and fundamental physical principles is essential. This limitation is particularly evident on challenging physical reasoning benchmarks, revealing a persistent gap in physical commonsense reasoning. To address this challenge, we propose PhysMRV, a training-free physical memory and verification framework for physical plausibility reasoning. Unlike retrieval-augmented VLMs that retrieve semantically similar videos as additional context, PhysMRV transforms training videos into a Hierarchical Memory Bank of structured physical knowledge comprising three complementary levels: scene descriptions capturing visual context, physical-event graphs modeling object interactions and causal structure, and physics-rule summaries distilling reusable physical principles and cues. During inference, PhysMRV retrieves physically relevant memories and leverages their structured physical evidence to guide a frozen VLM in verifying physical plausibility, requiring neither fine-tuning nor parameter updates. We evaluate PhysMRV on three challenging physical reasoning benchmarks, ImplausiBench, IntPhys2, and GRASP Level 2, across multiple state-of-the-art VLMs. Experimental results demonstrate consistent improvements over direct prompting across diverse VLMs and evaluation benchmarks, showing that structured physical memories provide an effective and scalable means of enhancing physical plausibility reasoning without additional training.
- Abstract(参考訳): ビデオ言語モデル(VLM)は、ビデオ理解と視覚的質問応答において顕著なパフォーマンスを達成しているが、オブジェクトの相互作用、因果ダイナミクス、基本的な物理原理を理解するという物理的妥当性の推論では信頼性が低いままである。
この制限は、物理的推論ベンチマークにおいて特に顕著であり、物理コモンセンス推論の持続的なギャップが明らかである。
この課題に対処するために、トレーニング不要な物理メモリと検証フレームワークであるPhysMRVを提案する。
意味論的に類似した動画を検索するVLMとは異なり、PhysMRVはトレーニングビデオを階層的な物理知識の階層的銀行に変換する。
推論中、PhysMRVは物理的に関連のある記憶を回収し、その構造化された物理的証拠を利用して、物理的妥当性を検証するために凍結されたVLMを誘導し、微調整もパラメータの更新も必要としない。
我々はPhysMRVを、複数の最先端VLMに対して、ImplausiBench、IntPhys2、GRASP Level 2の3つの挑戦的物理推論ベンチマークで評価した。
実験により,様々なVLMおよび評価ベンチマークの直接的プロンプトよりも一貫した改善が示された。
関連論文リスト
- NICE FACT: Diagnosing and Calibrating VLMs in Quantitative Reasoning for Kinematic Physics [65.02899948986969]
この研究は、視覚言語モデルが物理的世界をどのように知覚するかを根本的に理解し、物理法則を活用することを目的としている。
運動物理学の量的推論を明示的に分解する双対診断パラダイムであるNICEとFACTを提案する。
NICEは、我々の新しい地区インフォームドキャリブレーション手法と、信頼性の評価とキャリブレーションのための新しいメトリクスについて研究する。
論文 参考訳(メタデータ) (2026-05-08T20:17:44Z) - PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning [49.88366485306749]
現代のビデオ生成モデルは、視覚的にリアルなビデオを生成することができるが、物理法則に従わないことが多い。
本稿では,物理認識力を高めるため,映像生成モデルを導くための表現として,物理知識を捉えたPhysMasterを提案する。
論文 参考訳(メタデータ) (2025-10-15T17:59:59Z) - LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference [57.086932851733145]
ビデオ拡散モデルにおける直感的な物理を評価するトレーニング不要な方法であるLikePhysを紹介した。
現在のビデオ拡散モデルにおける直観的物理理解のベンチマークを行う。
経験的結果は、現在のモデルが複雑でカオス的な力学に苦しむにもかかわらず、モデルキャパシティと推論設定スケールとしての物理理解の改善傾向が明らかであることを示している。
論文 参考訳(メタデータ) (2025-10-13T15:19:07Z) - TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility [70.24211591214528]
ビデオ生成モデルは、浮動、テレポート、モーフィングのような直感的な物理法則に違反したシーケンスを生成する。
既存のビデオランゲージモデル(VLM)は、物理違反の特定に苦慮し、時間的および因果的推論における根本的な制限を明らかにしている。
我々は、バランスの取れたトレーニングデータセットと軌道認識型アテンションモジュールを組み合わせた微調整レシピTRAVLを導入し、モーションエンコーディングを改善する。
言語バイアスを除去し,視覚的時間的理解を分離する300本のビデオ(150本実写150本)のベンチマークであるImplausiBenchを提案する。
論文 参考訳(メタデータ) (2025-10-08T21:03:46Z) - PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding [21.91860938879665]
視覚言語モデル(VLM)は、常識的推論において優れているが、物理世界を理解するのに苦労していることを示す。
本稿では、VLMの一般化強度とビジョンモデルの専門知識を組み合わせたフレームワークであるPhysAgentを紹介する。
以上の結果から,VLMの物理世界理解能力の向上は,Mokaなどのエージェントの具体化に有効であることが示唆された。
論文 参考訳(メタデータ) (2025-01-27T18:59:58Z) - Physics Context Builders: A Modular Framework for Physical Reasoning in Vision-Language Models [11.282655911647483]
視覚言語モデル(VLM)における物理推論の課題
物理コンテキストビルダー(PCB)は,物理シーンの詳細な記述を生成するために,より小型のVLMを微調整したモジュラーフレームワークである。
PCBは、視覚知覚と推論の分離を可能にし、身体的理解に対する相対的な貢献を分析することができる。
論文 参考訳(メタデータ) (2024-12-11T18:40:16Z) - ContPhy: Continuum Physical Concept Learning and Reasoning from Videos [86.63174804149216]
ContPhyは、マシン物理常識を評価するための新しいベンチマークである。
私たちは、さまざまなAIモデルを評価し、ContPhyで満足なパフォーマンスを達成するのに依然として苦労していることがわかった。
また、近年の大規模言語モデルとパーティクルベースの物理力学モデルを組み合わせるためのオラクルモデル(ContPRO)を導入する。
論文 参考訳(メタデータ) (2024-02-09T01:09:21Z) - Dynamic Visual Reasoning by Learning Differentiable Physics Models from
Video and Language [92.7638697243969]
視覚概念を協調的に学習し,映像や言語から物体の物理モデルを推定する統合フレームワークを提案する。
これは視覚認識モジュール、概念学習モジュール、微分可能な物理エンジンの3つのコンポーネントをシームレスに統合することで実現される。
論文 参考訳(メタデータ) (2021-10-28T17:59:13Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。