論文の概要: What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning
- arxiv url: http://arxiv.org/abs/2607.16938v1
- Date: Sat, 18 Jul 2026 19:20:16 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-21 18:48:37.306378
- Title: What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning
- Title(参考訳): 複雑な道路シナリオを視覚-言語-行動モデルで解釈して、安全で信頼性の高い自動運転車の学習
- Authors: Kalpana Panda, Wesley Maia, Vinti Agarwal, Ross Greer,
- Abstract要約: フロントカメラ画像から検出対象を体系的に除去するフレームワークを提案する。
これにより、各オブジェクトの存在がモデルの計画行動に与える影響を分離する。
また、モデルが人間ドライバーが考慮する対象に強く反応するケースも見つかる。
- 参考スコア(独自算出の注目度): 2.5264159985739623
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: End-to-end autonomous driving models are now able to navigate complex road scenarios, mapping raw sensor observations directly to observed paths for open-loop evaluation and often effective driving in closed-loop evaluation. Yet the internal logic of these safety-critical systems remains largely opaque, due to the complexity of traffic scenes. We propose a counterfactual ablation framework called Counterfactual Vision Action Analysis (CVAA) that systematically removes individual detected objects from front-camera images using photorealistic generative inpainting to prepare counterfactual sets to evaluate the difference in the model's response. This isolates the causal effect of each object's presence on the model's planning behaviour. Applied to the Alpamayo 1 trajectory predictor across 210 nuScenes driving scenes, we create a dataset Counter -nuScenes, using which we see that vehicles and pedestrians within the model's 'path' dominate causal influence as expected, while traffic lights, as expected, exert disproportionate effect relative to their image footprint. However, we also find cases where the model responds strongly to objects a human driver would consider irrelevant. This brings forth a deeper question: does the model itself view the scene as a sum of individual objects influencing the outcome, or does it encode an entirely different set of internal features that do not correspond to human-legible scene elements? To further understand this, we compare intermediate representations of original and inpainted image pairs using mechanistic interpretability techniques and examine the effect of the removal through the various model layers. Together, these two stages offer a path from behavioral auditing to representational understanding, creating explainable driving systems and solidifying human-AI trust.
- Abstract(参考訳): エンド・ツー・エンドの自律運転モデルでは複雑な道路シナリオをナビゲートし、観測経路に直接生のセンサー観測をマッピングしてオープンループ評価を行い、クローズドループ評価においてしばしば効果的な運転を行うことができる。
しかし、これらの安全クリティカルなシステムの内部ロジックは、交通シーンの複雑さのため、ほとんど不透明である。
本稿では,画像から検出された個々の物体を,フォトリアリスティック・ジェネレーティブ・インペインティングを用いて系統的に除去し,モデル応答の差を評価できる,CVAA(Contererfactual Vision Action Analysis)という逆ファクト・アブレーション・フレームワークを提案する。
これにより、各オブジェクトの存在がモデルの計画行動に与える影響を分離する。
Alpamayo 1 トラジェクティブ予測器を210 nuScenes の走行シーンに適用し、データセットのカウンタ -nuScenes を作成し、モデル内の車や歩行者が期待通りに因果的影響を支配しているのに対して、交通信号は期待どおり、画像のフットプリントに対して不均等な効果を発揮することを確認します。
しかし、モデルが人間ドライバーが考慮する対象に強く反応するケースも見出す。
モデル自体が、結果に影響を与える個々のオブジェクトの総体としてシーンを捉えているのか、それとも、人間が適用可能なシーン要素に対応しない、まったく異なる内部的特徴のセットをエンコードしているのか?
これをさらに理解するために、機械的解釈可能性技術を用いてオリジナルとインペイントされた画像対の中間表現を比較し、様々なモデル層による除去の効果について検討する。
これら2つの段階は、行動監査から表現的理解へのパスを提供し、説明可能な運転システムを作成し、人間とAIの信頼を固める。
関連論文リスト
- Realistic pedestrian-driver interaction modelling using multi-agent RL with human perceptual-motor constraints [3.4087048085988765]
本稿では,歩行者エージェントとドライバエージェントの視覚的制約と運動的制約を統合したマルチエージェント強化学習フレームワークを提案する。
その結果,視覚的制約と運動的制約を併用したモデルが最適であることが示唆された。
本フレームワークは,人口レベルの分布として人的制約を制御するパラメータをモデル化することで,個人差を考慮に入れている。
論文 参考訳(メタデータ) (2025-10-31T11:18:13Z) - InsightDrive: Insight Scene Representation for End-to-End Autonomous Driving [3.8737986316149775]
我々はInsightDriveと呼ばれる新しいエンドツーエンドの自動運転手法を提案する。
言語誘導されたシーン表現によって知覚を整理する。
実験では、InsightDriveはエンドツーエンドの自動運転において最先端のパフォーマンスを達成する。
論文 参考訳(メタデータ) (2025-03-17T10:52:32Z) - QuAD: Query-based Interpretable Neural Motion Planning for Autonomous Driving [33.609780917199394]
自動運転車は環境を理解して適切な行動を決定する必要がある。
従来のシステムは、シーン内のエージェントを見つけるためにオブジェクト検出に依存していた。
我々は、最初に占有する時間的自律性を知覚するカスケードモジュールから遠ざかる、統一的で解釈可能で効率的な自律フレームワークを提案する。
論文 参考訳(メタデータ) (2024-04-01T21:11:43Z) - BEVSeg2TP: Surround View Camera Bird's-Eye-View Based Joint Vehicle
Segmentation and Ego Vehicle Trajectory Prediction [4.328789276903559]
軌道予測は自動車の自律性にとって重要な課題である。
学習に基づく軌道予測への関心が高まっている。
認識能力を向上させる可能性があることが示される。
論文 参考訳(メタデータ) (2023-12-20T15:02:37Z) - Linking vision and motion for self-supervised object-centric perception [16.821130222597155]
オブジェクト中心の表現は、自律運転アルゴリズムが多くの独立したエージェントとシーンの特徴の間の相互作用を推論することを可能にする。
伝統的にこれらの表現は教師付き学習によって得られてきたが、これは下流の駆動タスクからの認識を分離し、一般化を損なう可能性がある。
我々は、RGBビデオと車両のポーズを入力として、自己教師対象中心の視覚モデルを適用してオブジェクト分解を行う。
論文 参考訳(メタデータ) (2023-07-14T04:21:05Z) - Street-View Image Generation from a Bird's-Eye View Layout [95.36869800896335]
近年,Bird's-Eye View (BEV) の知覚が注目されている。
自動運転のためのデータ駆動シミュレーションは、最近の研究の焦点となっている。
本稿では,現実的かつ空間的に一貫した周辺画像を合成する条件生成モデルであるBEVGenを提案する。
論文 参考訳(メタデータ) (2023-01-11T18:39:34Z) - Exploring Contextual Representation and Multi-Modality for End-to-End
Autonomous Driving [58.879758550901364]
最近の知覚システムは、センサー融合による空間理解を高めるが、しばしば完全な環境コンテキストを欠いている。
我々は,3台のカメラを統合し,人間の視野をエミュレートするフレームワークを導入し,トップダウンのバードアイビューセマンティックデータと組み合わせて文脈表現を強化する。
提案手法は, オープンループ設定において0.67mの変位誤差を達成し, nuScenesデータセットでは6.9%の精度で現在の手法を上回っている。
論文 参考訳(メタデータ) (2022-10-13T05:56:20Z) - TRiPOD: Human Trajectory and Pose Dynamics Forecasting in the Wild [77.59069361196404]
TRiPODは、グラフの注目ネットワークに基づいて身体のダイナミクスを予測する新しい方法です。
実世界の課題を取り入れるために,各フレームで推定された身体関節が可視・視認可能かどうかを示す指標を学習する。
評価の結果,TRiPODは,各軌道に特化して設計され,予測タスクに特化している。
論文 参考訳(メタデータ) (2021-04-08T20:01:00Z) - Deep Structured Reactive Planning [94.92994828905984]
自動運転のための新しいデータ駆動型リアクティブ計画目標を提案する。
本モデルは,非常に複雑な操作を成功させる上で,非反応性変種よりも優れることを示す。
論文 参考訳(メタデータ) (2021-01-18T01:43:36Z) - Studying Person-Specific Pointing and Gaze Behavior for Multimodal
Referencing of Outside Objects from a Moving Vehicle [58.720142291102135]
物体選択と参照のための自動車応用において、手指しと目視が広く研究されている。
既存の車外参照手法は静的な状況に重点を置いているが、移動車両の状況は極めて動的であり、安全性に制約がある。
本研究では,外部オブジェクトを参照するタスクにおいて,各モダリティの具体的特徴とそれら間の相互作用について検討する。
論文 参考訳(メタデータ) (2020-09-23T14:56:19Z) - Implicit Latent Variable Model for Scene-Consistent Motion Forecasting [78.74510891099395]
本稿では,センサデータから直接複雑な都市交通のシーン一貫性のある動き予測を学習することを目的とする。
我々は、シーンを相互作用グラフとしてモデル化し、強力なグラフニューラルネットワークを用いてシーンの分散潜在表現を学習する。
論文 参考訳(メタデータ) (2020-07-23T14:31:25Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。