論文の概要: VepAgent: Bridging Causal-Transition via Tool-Augmented Reinforcement Learning for Video Event Prediction
- arxiv url: http://arxiv.org/abs/2610.06293v1
- Date: Mon, 05 Oct 2026 13:19:10 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-09 09:25:22.042917
- Title: VepAgent: Bridging Causal-Transition via Tool-Augmented Reinforcement Learning for Video Event Prediction
- Title(参考訳): VepAgent: ビデオイベント予測のためのツール強化強化学習による因果遷移のブリッジ
- Abstract要約: VepAgentは、因果遷移推論とツール強化強化学習を統合するエージェントフレームワークである。
提案手法は最先端の性能を達成し,より大きなMLLMよりも優れることを示す。
- 参考スコア(独自算出の注目度): 27.3895198744468
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in video understanding, yet their reliance on retrospective summarization and text-centric priors often limits their ability to bridge unobserved causal transitions when applied to Video Event Prediction (VEP). To address this, we propose VepAgent, an agentic framework that integrates causal-transition reasoning with tool-augmented reinforcement learning (RL) for robust VEP. Unlike prior methods that passively project future trajectories from historical dependencies, our approach explicitly models the logical progression from terminal observed states to future events. Specifically, we first construct futurebench-4K, a high-quality chain-of-thought dataset for supervised fine-tuning (SFT) that effectively bridges the causal-logic gap by structuring the deduction of unobserved intermediate states. Subsequently, we develop a diagnostic tool library integrating state tracking, frame retrieval, and region magnification, enabling the agent to dynamically augment reasoning with external tools to recover missing spatio-temporal evidence and resolve visual ambiguities during inference. Moreover, we propose a composite reward mechanism that jointly optimizes prediction accuracy, causal coherence, and reliable prior, compelling the agent to rely on genuine visual grounding rather than superficial textual similarities. Extensive evaluations on FutureBench and NEPBench datasets demonstrate that our method achieves state-of-the-art performance, significantly outperforming larger MLLMs and validating the empirical effectiveness of our agentic, future-oriented reasoning paradigm.
- Abstract(参考訳): MLLM(Multimodal Large Language Models)は、ビデオ理解において顕著な可能性を示しているが、振り返りの要約とテキスト中心の事前依存は、ビデオイベント予測(VEP)に適用した場合に、観測されない因果遷移をブリッジする能力を制限することがしばしばある。
そこで本研究では,因果遷移推論とツール強化強化学習(RL)を統合し,堅牢なVEPを実現するエージェントフレームワークであるVepAgentを提案する。
従来の手法と異なり,本手法では, 終末観測状態から将来の事象への論理的進行を明示的にモデル化する。
具体的には、まず、教師付き微調整(SFT)のための高品質なチェーン・オブ・シント・データセットであるFuturebench-4Kを構築し、未観測中間状態の推論を構造化することにより因果論理的ギャップを効果的に橋渡しする。
その後、状態追跡、フレーム検索、領域拡大を統合した診断ツールライブラリを開発し、外部ツールによる推論を動的に拡張し、欠落した時空間的証拠を回収し、推論中の視覚的曖昧性を解決する。
さらに,予測精度,因果コヒーレンス,信頼性の高い事前処理を共同で最適化する複合報酬機構を提案する。
FutureBench と NEPBench データセットの大規模評価により,本手法は最先端の性能を実現し,より大きなMLLM を著しく上回り,エージェント的,未来志向の推論パラダイムの実証的有効性を検証した。
関連論文リスト
- MEA: A Reward-Driven Multi-Agent System for Faithful Model Explanations [25.142295941569262]
我々は、説明知識の障壁を完全に取り除くマルチエージェントフレームワークMEAを提案する。
本稿では,特徴属性にまたがる多様な質問タイプ,反事実推論,突発的特徴検出について紹介する。
We found that Frontier LLMs produce systematicly of unfaithful explanations。
論文 参考訳(メタデータ) (2026-10-01T20:57:29Z) - Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs [68.58207076756237]
本稿では,結果評価からメカニズム診断へ移行する摂動に基づく評価プロトコルProCauEvalを紹介する。
因果推論において,ビデオコンテンツは体系的に過小評価されている。
教師のネガティブなアライメントに基づく強化学習フレームワークであるADPOを提案する。
論文 参考訳(メタデータ) (2026-05-10T08:48:58Z) - Large Vision-Language Models Get Lost in Attention [51.851592109135716]
本稿では,情報理論と幾何に基づく統合フレームワークを提案し,残差更新の幾何的およびエントロピー的性質を定量化する。
注意は再設定に焦点を当てたサブスペース言語演算子として機能し、FFNはセマンティックイノベーションを駆動するサブスペース言語演算子として機能します。
論文 参考訳(メタデータ) (2026-05-07T04:45:52Z) - Rethinking Video Human-Object Interaction: Set Prediction over Time for Unified Detection and Anticipation [18.693454975393703]
ビデオベースヒューマンオブジェクトインタラクション(HOI)の理解には、進行中のインタラクションを検出し、将来の進化を予測する必要がある。
対象対象の局所化,現在のHOI検出,将来の予測を共同で行う,ペア中心のフレームワークであるDETAnt-HOIとHOI-DAを紹介する。
実験では、検出と予測の両方において一貫した改善が見られ、より長い地平線でより大きな利得が得られた。
論文 参考訳(メタデータ) (2026-04-12T01:07:43Z) - Insight-V++: Towards Advanced Long-Chain Visual Reasoning with Multimodal Large Language Models [65.4947731385794]
基礎画像中心モデルであるInsight-Vから進化した統合多エージェント視覚推論フレームワークを提案する。
空間的時間的推論を強化し、評価ロバスト性を向上させる2つの新しいアルゴリズムST-GRPOとJ-GRPOを導入する。
LLaVA-NeXTやQwen2.5-VLといったベースモデルの実験は、挑戦的な画像とビデオの推論ベンチマーク間で大きなパフォーマンス向上を示している。
論文 参考訳(メタデータ) (2026-03-18T15:28:07Z) - From Passive Metric to Active Signal: The Evolving Role of Uncertainty Quantification in Large Language Models [77.04403907729738]
このサーベイは、受動的診断基準からリアルタイムモデル動作を導くアクティブ制御信号への不確実性の進化をグラフ化する。
3つのフロンティアにまたがるアクティブ制御信号として不確実性がいかに活用されているかを示す。
この調査は、次世代のスケーラブルで信頼性があり、信頼できるAIを構築するためには、新しい不確実性のトレンドを習得することが不可欠である、と論じている。
論文 参考訳(メタデータ) (2026-01-22T06:21:31Z) - Current Agents Fail to Leverage World Model as Tool for Foresight [61.82522354207919]
エージェントは、行動する前に結果を予測するためにそれらを使用できます。
本稿では,現在のエージェントがそのような世界モデルを,認知力を高めるツールとして活用できるかどうかを実証的に検討する。
論文 参考訳(メタデータ) (2026-01-07T13:15:23Z) - Beyond Patterns: Harnessing Causal Logic for Autonomous Driving Trajectory Prediction [10.21659221112514]
本稿では、因果推論を利用して予測堅牢性、一般化、精度を向上させる新しい軌道予測フレームワークを提案する。
本研究は、軌跡予測の因果推論の可能性を強調し、ロバストな自律運転システムへの道を開くものである。
論文 参考訳(メタデータ) (2025-05-11T05:56:07Z) - Unified Recurrence Modeling for Video Action Anticipation [16.240254363118016]
本稿では,メッセージパッシングフレームワークを用いたビデオアクション予測のための統合再帰モデルを提案する。
提案手法は,EPIC-Kitchenデータセットの大規模化において,従来よりも優れている。
論文 参考訳(メタデータ) (2022-06-02T12:16:44Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。