論文の概要: CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention
- arxiv url: http://arxiv.org/abs/2608.04396v1
- Date: Wed, 05 Aug 2026 02:58:41 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-06 14:48:43.698323
- Title: CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention
- Title(参考訳): CofactVLA: 対物干渉によるビジョン・ランゲージ・アクションモデル
- Authors: Yan Zhang, Yinan Wu, Haoran Duan, Jungong Han,
- Abstract要約: VLA(Vision-Language-Action)モデルはロボット操作に大きな進歩をもたらしたが、ビジョンオーバーライド現象に苦戦している。
本稿では,新たな因果介入フレームワークCofactVLAを提案する。
我々は、CofactVLAが様々なシミュレーションベンチマークにまたがって新しい最先端のベンチマークを確立していることを示す。
- 参考スコア(独自算出の注目度): 48.20065816234311
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Vision-Language-Action (VLA) models have driven significant progress in robotic manipulation, yet they fundamentally struggle with the vision-override phenomenon. Driven by the severe modality imbalance between dense visual streams and sparse linguistic instructions, VLAs frequently fall prey to causal confusion. Instead of treating language as the primary causal driver, the policy entirely bypasses the original instruction by overfitting to spurious visual confounders, such as prominent objects or familiar layouts. To systematically alleviate this bias, we formalize the process of action generation as a Dual-path Deconfounding Graph (DDG) and propose CofactVLA, a novel causal intervention framework. By dynamically constructing a language-masked counterfactual branch within a single forward pass, CofactVLA isolates and neutralizes visual confounders through two synergistic mechanisms. First, Action-Level Orthogonal Projection Guidance (OPG) geometrically projects the factual velocity field away from the counterfactual visual bias during continuous flow matching, extracting the pure semantic intent. Second, Feature-Level Counterfactual Covariance Reduction (CCR) mathematically deconfounds latent representations by penalizing the positive eigenspace of the covariance difference, explicitly suppressing dominant visual shortcuts while preserving the causal language intent. Extensive experiments demonstrate that CofactVLA establishes a new state-of-the-art across diverse simulation benchmarks. Beyond simulation, real-world robot experiments demonstrate the causal efficacy of our method in bridging the generalization gap, yielding a 52.3\% absolute success rate gain under out-of-distribution scenarios.
- Abstract(参考訳): VLA(Vision-Language-Action)モデルはロボット操作に大きな進歩をもたらしたが、基本的にビジョンオーバーライド現象に苦戦している。
濃密な視覚ストリームと希薄な言語指示の間の重大なモダリティの不均衡によって、VLAは因果的混乱に陥ることが多い。
言語を主要な因果的ドライバとして扱う代わりに、このポリシーは、著名なオブジェクトや慣れ親しんだレイアウトなど、刺激的な視覚的共同創設者に過度に適合させることによって、オリジナルの命令を完全に回避する。
このバイアスをシステマティックに緩和するために、Dual-path Deconfounding Graph (DDG)としてアクション生成のプロセスを形式化し、新しい因果介入フレームワークであるCofactVLAを提案する。
CofactVLAは、1つのフォワードパス内で言語にマッチした反ファクトのブランチを動的に構築することにより、2つのシナジスティックメカニズムを通じて視覚的共同創設者を分離し、中和する。
第一に、アクション・レベル直交射影誘導(OPG)は、連続フローマッチング中の反ファクト的視覚バイアスから実際の速度場を幾何学的に投影し、純粋な意味的意図を抽出する。
第二に、特徴レベル対事実共分散低減(CCR)は、共分散差の正の固有空間をペナルティ化し、因果言語意図を保ちながら支配的な視覚的ショートカットを明示的に抑制することにより、潜在表現を数学的に分解する。
大規模な実験により、CofactVLAは多様なシミュレーションベンチマークにまたがって新しい最先端のベンチマークを確立している。
シミュレーションの他に、実世界のロボット実験は、一般化ギャップを埋める際の方法の因果効果を実証し、アウト・オブ・ディストリビューションシナリオ下では52.3倍の絶対成功率を得る。
関連論文リスト
- CV-DCLR: Causal-Visual Dynamic Label Refinement for Robust Zero-Shot Learning [25.21342780360303]
本稿ではCausal-Visual Dynamic Label Refinementフレームワークを提案する。
二重ストリーム相互補正機構を用いて視覚的意味的関連を補正する。
CUB、SUN、AWA2ベンチマークの実験では、CV-DCLRは高いあいまいさのシナリオで最先端の手法を大幅に上回っている。
論文 参考訳(メタデータ) (2026-07-01T12:45:35Z) - ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation [71.69632267962993]
我々は、操作計画が自然に予測に分解され、次の視覚状態が予測され、逆ダイナミクスとなることを論じる。
我々は、この分解を実現する生成モデルである textbfThinkingVLA を提案する。
シミュレーションと実世界のベンチマークの実験では、ThinkingVLAは最先端のベースラインを一貫して上回っている。
論文 参考訳(メタデータ) (2026-06-16T13:45:17Z) - Attention-space Contrastive Guidance for Efficient Hallucination Mitigation in LVLMs [9.043999205886658]
大きな視覚言語モデルにおける幻覚は、言語が視覚的証拠を支配するときにしばしば起こる。
本稿では,視覚言語と言語のみの注意経路を構築するために,自己注意層内で動作するシングルパス機構であるContrastive Guidance(ACG)を提案する。
ACGは、計算コストを大幅に削減しつつ、最先端の忠実さとキャプション品質を達成する。
論文 参考訳(メタデータ) (2026-01-20T08:04:18Z) - Stable Language Guidance for Vision-Language-Action Models [62.80963701282789]
残留セマンティックステアリング(Residual Semantic Steering)は、セマンティック実行から身体的余裕を逸脱する確率的フレームワークである。
RSSは最先端の堅牢性を実現し、敵対的な言語摂動の下でも性能を維持する。
論文 参考訳(メタデータ) (2026-01-07T16:16:10Z) - When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models [75.16145284285456]
我々は,White-box設定とBlack-box設定の両方の下で,組込みVLAモデルのマルチモーダル対向ロバスト性に関する総合的研究であるVLA-Foolを紹介する。
自動生成および意味的に誘導されるプロンプトフレームワークを最初に開発する。
LIBEROベンチマークの実験では、小さなマルチモーダル摂動でさえ大きな行動偏差を引き起こすことが示されている。
論文 参考訳(メタデータ) (2025-11-20T10:14:32Z) - HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model [54.64088247291416]
操作ポリシー設計の基本的な目的は、ロボットに人間の指示を理解し、シーンの手がかりを推論し、動的な環境で一般化されたアクションを実行することである。
近年の自己回帰的視覚言語行動(VLA)法は、視覚言語モデル(VLM)から常識推論能力を継承し、次の行動予測を行う。
拡散に基づく行動の連続的な性質と自己回帰の文脈的推論を吸収する統合フレームワークであるHybridVLAを紹介する。
論文 参考訳(メタデータ) (2025-03-13T17:59:52Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。