論文の概要: DS-VLA: A Dendritic-inspired Vision-Language-Action Model for Robust Action Control
- arxiv url: http://arxiv.org/abs/2609.32253v1
- Date: Sat, 26 Sep 2026 05:17:16 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-07 18:28:28.622097
- Title: DS-VLA: A Dendritic-inspired Vision-Language-Action Model for Robust Action Control
- Title(参考訳): DS-VLA:ロバスト動作制御のための樹状刺激型視覚言語行動モデル
- Abstract要約: DS-VLAは樹状体にインスパイアされたアクションアーキテクチャであり、樹状体スパイキングダイナミクスをVLA制御に組み込む。
我々は4つのLIBEROスイートのDS-VLAを、名目上ロールアウトと統一されたクローズドループ動作摂動プロトコルの両方で評価する。
- 参考スコア(独自算出の注目度): 15.462788061621572
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Vision-language-action (VLA) models have achieved strong performance in language-conditioned manipulation, yet success under nominal evaluation does not necessarily translate into robust closed-loop behavior when executed actions are transiently corrupted. We introduce DS-VLA, a dendritic-inspired action architecture that incorporates dendritic spiking dynamics into VLA control to address this limitation. Specifically, to enable modularized feature processing and temporal information integration, DS-VLA equips action neurons with multiple sparsely connected dendritic branches, each featuring heterogeneous, learned decay factors. Furthermore, to suppress unreliable state updates while preserving task-relevant historical information, we introduce a neuron-wise inhibitory gate that adaptively regulates the admission of new multimodal evidence into dendritic states prior to somatic dynamics. We evaluate DS-VLA on all four LIBERO suites under both nominal rollouts and a unified closed-loop action-perturbation protocol. DS-VLA achieves a 91.6\% average nominal success rate and an 87.35\% average perturbed success rate, retaining 95.4\% of its nominal performance. Under the same reported perturbation setting, OpenVLA-OFT, FAST, $π_0$, and GR00T achieve 39.45\%, 23.90\%, 28.55\%, and 30.75\%, respectively. A controlled ablation isolates the contribution of neuron-wise shared inhibition, while analyses of neural dynamics and post-perturbation trajectories associate robust performance with selective evidence suppression and effective behavioral recovery. Together, these results demonstrate that integrating brain-inspired computational mechanisms offers a promising architectural prior for robust embodied intelligence beyond merely scaling vision-language backbones or generative action decoders.
- Abstract(参考訳): VLA(Vision-Language-action)モデルは、言語条件の操作において高い性能を達成しているが、名目評価による成功は、実行された動作が過渡的に破損した場合に、必ずしも堅牢な閉ループ動作に変換されない。
DS-VLAは樹状体にインスパイアされたアクションアーキテクチャで, 樹状体のスパイクダイナミクスをVLA制御に組み込んで, この制限に対処する。
具体的には、モジュール化された特徴処理と時間情報統合を可能にするため、DS-VLAは複数の疎結合な樹状突起枝を持つ作用ニューロンをそれぞれ不均一で学習された崩壊因子を特徴とする。
さらに,タスク関連履歴情報を保存しながら,信頼性の低い状態更新を抑えるため,身体力学に先立って,新しいマルチモーダル証拠の樹状状態への入力を適応的に規制するニューロン抑制ゲートを導入する。
我々は4つのLIBEROスイートのDS-VLAを、名目上ロールアウトと統一されたクローズドループ動作摂動プロトコルの両方で評価する。
DS-VLAは、平均的な名目成功率91.6\%、平均的な摂動成功率87.35\%を達成し、名目パフォーマンスの95.4\%を維持している。
同じ報告された摂動条件の下で、OpenVLA-OFT、FAST、$π_0$、GR00Tはそれぞれ39.45\%、23.90\%、28.55\%、30.75\%となる。
制御されたアブレーションは、神経学的に共有された抑制の寄与を分離する一方、神経力学と摂動後軌道の分析は、堅牢な性能と選択的エビデンス抑制と効果的な行動回復とを関連付ける。
これらの結果は、脳にインスパイアされた計算機構の統合が、単に視覚言語バックボーンや生成アクションデコーダをスケーリングするだけでなく、堅牢なインボディードインテリジェンスに有望なアーキテクチャ的事前を提供することを示した。
関連論文リスト
- CereVLA: Cerebellum-Inspired Consequence-Aware Residual Governance for Efficient Vision-Language-Action Execution [53.635273213001796]
アクションチャンクされた視覚言語アクション(VLA)ポリシーは推論効率を改善するが、コミットされたアクションチャンク内の限られたフィードバックは、実行エラーの蓄積につながる。
本稿では,軽量残効改善と予測結果評価を冷凍VLA実行に統合した統合フレームワークであるCereVLAについて述べる。
論文 参考訳(メタデータ) (2026-09-23T07:30:58Z) - TrustVLA: Mechanism-Guided Inference-Time Defense Against Vision-Language-Action Backdoors [54.12833308568015]
VLA(Vision-Language-Action)モデルは、エンドユーザが監査できないパイプラインを通じてデプロイされる。
小さな視覚トリガーは、障害が観測可能である前に、長い水平ロボットポリシーをリダイレクトする。
個別に提案した2つのVLA攻撃を通してこのギャップについて検討した。
論文 参考訳(メタデータ) (2026-07-14T09:43:52Z) - RotVLA: Rotational Latent Action for Vision-Language-Action Model [54.22746299071677]
本稿では,連続的な回転潜在動作表現に基づくVLAフレームワークであるRotVLAを紹介する。
潜在作用はSO(n) の元としてモデル化され、連続性、構成性、および実世界の作用力学と整合した構造的幾何学を提供する。
RotVLAはVLMバックボーンとフローマッチングアクションヘッドで構成される。
論文 参考訳(メタデータ) (2026-05-13T11:58:02Z) - Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery [62.75419724651416]
textbfSentinel-VLAは,リアルタイム実行状況を監視するアクティブセンチネルモジュールを備えたメタ認知型VLAモデルである。
すべてのトレーニングデータは、設計したパイプラインを通じて自動生成され、注釈付けされます。
実世界の実験では、Sentinel-VLAはSOTAモデルであるPI0と比較してタスク成功率を30%以上向上することを示した。
論文 参考訳(メタデータ) (2026-05-02T02:10:54Z) - HELM: Harness-Enhanced Long-horizon Memory for Vision-Language-Action Manipulation [2.9231828959903474]
VLA(Vision-Language-Action)モデルは、短軸性能が強いにもかかわらず、長軸操作タスクにおいて体系的に失敗する。
この失敗は、現在のリアクティブ実行設定でコンテキスト長だけを拡張することで解決されないことを示す。
HELMは3つのコンポーネントでこれらの欠陥に対処するモデルに依存しないフレームワークである。
論文 参考訳(メタデータ) (2026-04-20T19:57:35Z) - STRONG-VLA: Decoupled Robustness Learning for Vision-Language-Action Models under Multimodal Perturbations [26.063335767640083]
本稿では、VLA(Vision-Language-Action)モデルのための切り離された微調整フレームワークSTRONG-VLAを提案する。
ステージIでは、モデルは困難が増す多モーダル摂動のカリキュラムに晒される。
ステージIIでは、モデルはクリーンなタスク分布と整合して、堅牢性を維持しながら実行の忠実さを回復します。
LIBEROベンチマークの実験では、STRONG-VLAは複数のVLAアーキテクチャにおけるタスク成功率を一貫して改善している。
論文 参考訳(メタデータ) (2026-04-11T06:37:47Z) - Self-Correcting VLA: Online Action Refinement via Sparse World Imagination [55.982504915794514]
本稿では, 自己補正VLA (SC-VLA) を提案する。
SC-VLAは最先端のパフォーマンスを達成し、最高タスクスループットを16%削減し、最高パフォーマンスのベースラインよりも9%高い成功率を得る。
論文 参考訳(メタデータ) (2026-02-25T06:58:06Z) - Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation [95.89924101984566]
GPM(Global Prior Memory)とLCM(Local Consistency Memory)を備えたデュアルメモリVLAフレームワークOptimusVLAを紹介する。
GPMはガウスノイズを意味論的に類似した軌道から取得したタスクレベルの先行値に置き換える。
LCMは、時間的コヒーレンスと軌道の滑らかさを強制する学習された一貫性制約を注入する。
論文 参考訳(メタデータ) (2026-02-22T15:39:34Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。