論文の概要: BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation
- arxiv url: http://arxiv.org/abs/2608.05042v1
- Date: Wed, 05 Aug 2026 16:54:25 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-06 14:48:44.015854
- Title: BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation
- Title(参考訳): BridgeVLA++: 3D操作のためのデータ効率が高く、一般化可能で、メモリ拡張されたビジョンランゲージ・アクションフレームワーク
- Authors: Peiyan Li, Yuze Zhu, Yixiang Chen, Qisen Ma, Yuan Xu, Jiabing Yang, He Guan, Yan Huang, Hongtao Wu, Xiao Ma, Tao Kong, Liang Wang, Tieniu Tan,
- Abstract要約: 視覚言語モデル(VLA)を構築するために、事前訓練された視覚言語モデル(VLM)を活用することが、3Dロボット操作の有望なパラダイムとして浮上している。
既存の3D VLA法はデータ不足であり、分布シフトの下での限定的な一般化を示し、過去の観測の明示的な記憶を欠いている。
持続的空間コンテキストと時間的相互作用履歴をモデル化した統合メモリアーキテクチャを用いて,BridgeVLA++を開発した。
- 参考スコア(独自算出の注目度): 40.910257242878956
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, exhibit limited generalization under distribution shifts, and lack explicit memory of past observations. These limitations hinder their application to data-scarce, open-world, and memory-dependent manipulation scenarios. Our previous work, BridgeVLA, improves data efficiency and generalization by preserving the input--output alignment of a pre-trained VLM during 3D action learning: raw point clouds are projected into multi-view images, and intermediate heatmaps are predicted before generating robot actions. In this work, we develop BridgeVLA++ by equipping BridgeVLA with a unified spatio-temporal memory architecture that models persistent spatial context and temporal interaction history. The resulting memory-augmented framework can reason over observation histories while preserving BridgeVLA's data efficiency and generalization capabilities. Extensive experiments show that our framework achieves strong performance on spatial manipulation tasks while exhibiting robust generalization. BridgeVLA++ further achieves state-of-the-art performance on two challenging memory-dependent manipulation benchmarks without sacrificing the data efficiency and generalization of the original BridgeVLA. In addition, BridgeVLA++ performs effectively in bimanual manipulation settings and is validated on an additional real-world robotic platform, demonstrating its scalability across tasks, environments, and robotic platforms. These results establish BridgeVLA++ as a unified 3D vision-language-action framework that simultaneously supports data-efficient learning, robust generalization, and effective memory-aware robot manipulation. Project website: https://bridgevla-plus.github.io/.
- Abstract(参考訳): 視覚言語モデル(VLA)を構築するために、事前訓練された視覚言語モデル(VLM)を活用することが、3Dロボット操作の有望なパラダイムとして浮上している。
しかし、既存の3D VLA法はデータ不足のままであり、分布シフトの下での限定的な一般化を示し、過去の観測の明示的な記憶を欠いている。
これらの制限は、データスカース、オープンワールド、メモリ依存の操作シナリオへのアプリケーションを妨げます。
これまでの研究であるBridgeVLAは,3次元動作学習において,事前学習したVLMの入力出力アライメントを保存することにより,データの効率と一般化を改善する。
本研究では,BridgeVLAを空間的コンテキストと時間的相互作用履歴をモデル化した一貫した時空間メモリアーキテクチャで実装することにより,BridgeVLA++を開発する。
結果として生じるメモリ拡張フレームワークは、BridgeVLAのデータ効率と一般化能力を保ちながら、観測履歴を解析することができる。
大規模な実験により,我々のフレームワークは,頑健な一般化を図りながら,空間操作タスクにおいて高い性能を発揮することが示された。
BridgeVLA++は、元のBridgeVLAのデータ効率と一般化を犠牲にすることなく、2つの挑戦的なメモリ依存の操作ベンチマークで最先端のパフォーマンスを実現する。
さらに、BridgeVLA++は、双方向操作設定で効果的に動作し、タスク、環境、ロボットプラットフォーム間のスケーラビリティを実証する、追加の現実世界のロボットプラットフォームで検証されている。
これらの結果は、データ効率の学習、堅牢な一般化、効果的なメモリ認識ロボット操作を同時にサポートする、統合された3Dビジョン言語アクションフレームワークとして、BridgeVLA++を確立する。
プロジェクトウェブサイト:https://bridgevla-plus.github.io/.com
関連論文リスト
- ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning [31.000965640377128]
ABot-M0は、システマティックデータキュレーションパイプラインを構築するフレームワークである。
これは不均一な生データを統一的で効率的な表現にエンドツーエンドに変換することを可能にする。
ABot-M0はデュアルストリーム機構を通じてモジュール認識をサポートする。
論文 参考訳(メタデータ) (2026-02-11T16:47:01Z) - GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning [20.646039344274556]
GeneralVLAは階層型視覚言語アクション(VLA)モデルであり、基礎モデルの一般化をより効果的に活用することができる。
GeneralVLAは14タスクの軌道生成に成功し、VoxPoserのような最先端の手法を著しく上回った。
論文 参考訳(メタデータ) (2026-02-04T08:30:27Z) - BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models [37.699828966838986]
BridgeVLAは、3D入力を複数の2D画像に投影し、VLMバックボーンとの入力アライメントを保証する新しい3D VLAモデルである。
アクション予測に2Dヒートマップを使用し、一貫した2次元画像空間内の入力空間と出力空間を統一する。
10以上のタスクで96.8%の成功率を達成することができ、1タスクにつき3つの軌道しか持たず、異常なサンプル効率を誇示している。
論文 参考訳(メタデータ) (2025-06-09T17:36:34Z) - HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation [54.03004125910057]
階層型視覚-言語-アクションモデルは、標準的なモノリシックVLAモデルよりも、ドメイン外のデータを利用するのに効果的であることを示す。
階層設計により、高レベルなVLMは、オフドメイン微調整データと実ロボットテストシナリオの間の重要なドメインギャップをまたいで転送可能であることを示す。
論文 参考訳(メタデータ) (2025-02-08T07:50:22Z) - TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies [95.30717188630432]
VLAモデルの行動予測のための時空間認識を容易にするために,視覚的トレースプロンプトを導入する。
我々は,これまでに収集した150Kロボット操作トラジェクトリのデータセットに基づいてOpenVLAを微調整し,新しいTraceVLAモデルを開発した。
4B Phi-3-Vision に基づくコンパクトな VLA モデルを提案する。
論文 参考訳(メタデータ) (2024-12-13T18:40:51Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。