論文の概要: CausalDriveBench: Evaluating Causal Reasoning in Vision-Language-Action Models for Autonomous Driving
- arxiv url: http://arxiv.org/abs/2609.32157v1
- Date: Sat, 26 Sep 2026 02:21:05 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-07 19:27:25.235838
- Title: CausalDriveBench: Evaluating Causal Reasoning in Vision-Language-Action Models for Autonomous Driving
- Title(参考訳): CausalDriveBench: 自律運転のための視覚言語行動モデルにおける因果推論の評価
- Abstract要約: CausalDriveBenchはPearlのCausal Hierarchy(PCH)に基盤を置く評価フレームワークである
我々は、因果的活動、休眠、気晴らしを区別する nuScene 上の因果的シーングラフを構築した。
ベンチマークには7,285の因果QAペアと1,000の反事実軌跡が含まれている。
- 参考スコア(独自算出の注目度): 9.843010103670167
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Vision-Language-Action (VLA) models for autonomous driving produce natural-language reasoning alongside predicted trajectories, but whether this reasoning reflects the causal structure of the scene remains untested. We introduce CausalDriveBench, an evaluation framework grounded in Pearl's Causal Hierarchy (PCH) that tests causal reasoning in driving-specific VLAs through structured visual question answering (QA) and alternative-trajectory prediction. To this end, we construct causal scene graphs over nuScenes that distinguish causally active, dormant, and distractor entities, separating perceptual salience from causal relevance. The benchmark spans all four rungs of PCH (association, intervention, and counterfactual along with causal discovery) for QA generation. For the higher rungs, we additionally provide reference trajectories under specified scene modifications, enabling action-level verification that complements reasoning-level evaluation. In total, the benchmark contains 7,285 verified causal QA pairs and 1,000 counterfactual trajectories derived from nuScenes. We evaluate 10 driving-specific VLAs and 3 general-purpose VLMs, and report three findings. First, the best model reaches only 70.6% QA accuracy, and 4 of 13 models score below random chance. Second, comparing each driving VLA to the general-purpose VLM that shares its language backbone, the cost of driving fine-tuning ranges from 2 to 34 percentage points on causal QA, with post-training design explaining the spread. Third, causal QA and trajectory accuracy are statistically uncorrelated across models: under counterfactual prompts, predicted trajectories either over-react or collapse onto the observed-scene baseline. Taken together, these results show that neither fluent rationales nor accurate observed-scene trajectories constitute evidence of causal understanding.
- Abstract(参考訳): 自律走行のためのVLA(Vision-Language-Action)モデルは、予測軌跡とともに自然言語推論を生成するが、この推論がシーンの因果構造を反映しているかどうかはまだ検証されていない。
本稿では,Pearl's Causal Hierarchy (PCH) に基礎を置く評価フレームワークCausalDriveBenchを紹介する。
この目的のために、我々は因果関係から知覚的サリエンスを分離し、因果関係、休息、気晴らしを区別する nuScenes 上の因果シーングラフを構築した。
このベンチマークは、QA生成のための4つのPCH(連想、介入、反ファクト)すべてにまたがる。
より高次ラングに対しては、特定のシーン修正の下で参照トラジェクトリも提供し、推論レベルの評価を補完するアクションレベルの検証を可能にする。
ベンチマークには7,285の因果QAペアと、nuScenesから派生した1,000の反事実軌道が含まれている。
運転特定VLA10例,汎用VLM3例について検討し,3例を報告する。
第一に、最良のモデルは70.6%のQA精度にしか達せず、13モデルのうち4つはランダムな確率以下である。
第二に、各駆動VLAを、言語バックボーンを共有する汎用VLMと比較すると、微調整のコストは、因果QAにおいて2~34ポイントであり、その拡散を説明する後トレーニング設計である。
第3に、因果的QAと軌跡的精度は、反ファクト的プロンプトの下では、観測時ベースラインに過度に反応するか、崩壊するかの予測軌跡を統計的に相関しない。
これらの結果は, 流理的理性や正確な観察・場面の軌跡が因果的理解の証拠となり得ないことを示唆している。
関連論文リスト
- Who Drives the Probability Game of VLMs? A Temporal Causal Drive Evaluation Framework [3.671673502935064]
本稿では,自己回帰復号における視覚入力,質問文,生成プレフィックスの役割を追及する因果的・時間的評価フレームワークを提案する。
MAVIS、LLaVA-Video-178K、MiraDataのQwen3-VL-8Bインストラクションの実験は、InternVL2-8Bのクロスモデル検証とともに、より強い早期質問からの一貫した遷移と、生成されたプレフィックスへの依存を高めるための視覚的ガイダンスを明らかにした。
論文 参考訳(メタデータ) (2026-09-02T02:19:36Z) - CAViAR: A Causal Video Dataset for Fine-Grained Accident Reasoning in Real-World Scenarios [10.195587813770075]
実世界の事故ビデオ2,249本からなる人手によるダッシュカムベンチマークであるCAViAR(Causal Accident Video and Incident Analysis Repository)を紹介する。
各ビデオには、環境条件、事故タイプ、因果説明、明らかなAt-Fault Agent、影響を受けたエージェント、および明らかなルール違反のカテゴリを含む構造化ラベルが注釈付けされている。
我々は、Cosmos-Reason2、Qwen3-VL、InternVL3を含む最先端のビジョン言語モデル(VLM)をベンチマークする。
論文 参考訳(メタデータ) (2026-08-19T18:51:21Z) - From Correlation to Causation in Lane Change Prediction for Automated Driving: A Causal Explanation Framework [11.81933456866472]
車線変更予測はインテリジェントな車両における中心的な課題である。
本稿では,車線変化予測と説明のための因果推論に基づくフレームワークを提案する。
論文 参考訳(メタデータ) (2026-06-14T11:32:20Z) - Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning [71.19675094463834]
この作業では、モデルが実行前に計画されたアクションを推論し、修正することを可能にする、自己修正型のVLAフレームワークである、Counterfactual VLAを導入している。
CF-VLAはまず、駆動意図を要約した時間分割メタアクションを生成し、その後、メタアクションと視覚コンテキストの両方で条件付けられた反実的推論を実行する。
大規模運転データセットの実験では、CF-VLAは軌道精度を最大17.6%向上し、安全基準を20.5%向上し、適応的思考を示す。
論文 参考訳(メタデータ) (2025-12-30T19:04:17Z) - Compressed Causal Reasoning: Quantization and GraphRAG Effects on Interventional and Counterfactual Accuracy [0.0]
本研究は, パールズ・コーサル・ラダーの全3レベルにわたる定量化効果を系統的に評価した。
Llama 3 8Bのラングレベルの精度は、量子化下では広く安定であり、NF4は全体の1%未満の劣化を示した。
CRASSベンチマークの実験では、既存のコモンセンスの反事実データセットには、量子化による推論ドリフトを明らかにするのに必要な構造感度が欠如していることが示されている。
論文 参考訳(メタデータ) (2025-12-13T17:54:15Z) - Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail [85.47497935739936]
Alpamayo-R1 (AR1) は、因果推論の連鎖と軌道計画を統合する視覚言語モデルである。
また,AR1は,軌道のみのベースラインに比べて,難問の計画精度が12%向上することを示した。
今後のアップデートで、AR1モデルとCoCのサブセットをリリースする予定です。
論文 参考訳(メタデータ) (2025-10-30T01:25:34Z) - SEAL: Steerable Reasoning Calibration of Large Language Models for Free [58.931194824519935]
大規模言語モデル(LLM)は、拡張チェーン・オブ・ソート(CoT)推論機構を通じて複雑な推論タスクに魅力的な機能を示した。
最近の研究では、CoT推論トレースにかなりの冗長性が示されており、これはモデル性能に悪影響を及ぼす。
我々は,CoTプロセスをシームレスに校正し,高い効率性を示しながら精度を向上する,トレーニング不要なアプローチであるSEALを紹介した。
論文 参考訳(メタデータ) (2025-04-07T02:42:07Z) - CausalVAE: Structured Causal Disentanglement in Variational Autoencoder [52.139696854386976]
変分オートエンコーダ(VAE)の枠組みは、観測から独立した因子をアンタングルするために一般的に用いられる。
本稿では, 因果内因性因子を因果内因性因子に変換する因果層を含むVOEベースの新しいフレームワークCausalVAEを提案する。
その結果、CausalVAEが学習した因果表現は意味論的に解釈可能であり、DAG(Directed Acyclic Graph)としての因果関係は精度良く同定された。
論文 参考訳(メタデータ) (2020-04-18T20:09:34Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。