論文の概要: Unleashing More Actions via Action Compositional Training for VLA Models
- arxiv url: http://arxiv.org/abs/2607.00351v1
- Date: Wed, 01 Jul 2026 02:48:17 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-02 19:56:07.693385
- Title: Unleashing More Actions via Action Compositional Training for VLA Models
- Title(参考訳): VLAモデルのための行動構成訓練によるより多くの行動の解放
- Authors: Kai Peng, Jie Lu, Xiaojiang Peng,
- Abstract要約: Vision-Language-Actionモデルは、デモデータのスケールと多様性によって駆動されるロボット操作において優れている。
オフラインデータ拡張フレームワークであるACT-VLA(Action Compositional Training for VLA Models)を提案する。
そこで本手法では,手作業によるデータ収集を不要にすることで,トレーニング分布を自動的に拡張し,過度な適合を緩和する。
- 参考スコア(独自算出の注目度): 23.837636418111035
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Vision-Language-Action models excel at robotic manipulation, driven by the scale and diversity of demonstration data. However, standard training paradigms often cause VLA models to severely overfit to specific behavioral patterns, rendering them unable to generalize to out-of-distribution scenarios even when those scenarios merely require novel combinations of identical sub-skills. While expanding datasets can mitigate this overfitting, acquiring high-quality robot data remains notoriously labor-intensive and cost-prohibitive. To resolve this impasse without expensive human teleoperation and to truly unleash more actions,i.e., enable VLA models to compose known sub-skills into a much broader set of executable behaviors beyond the original demonstrations-we propose ACT-VLA (Action Compositional Training for VLA Models), an offline data augmentation framework that leverages the model's latent task representations to synthesize novel, physically valid demonstrations directly from existing tasks for policy training. By eliminating additional manual data collection, our method automatically expands the training distribution and mitigates overfitting. We evaluate our approach on challenging manipulation tasks in simulation. Experiments demonstrate that while baseline VLA models generalize poorly due to original distribution overfitting, policies trained with our synthesized data achieve substantially higher success rates, validating that leveraging existing tasks for automated demonstration synthesis provides an effective, scalable, and data-efficient route to broadening VLA generalization.
- Abstract(参考訳): Vision-Language-Actionモデルは、デモデータのスケールと多様性によって駆動されるロボット操作において優れている。
しかしながら、標準的な訓練パラダイムは、VLAモデルを特定の行動パターンに強く適合させ、それらのシナリオが単に同一のサブスキルの新たな組み合わせを必要とする場合であっても、配布外シナリオに一般化できないようにする。
データセットを拡大すれば、この過度な適合を緩和できるが、高品質なロボットデータを取得することは、労働集約的でコストを抑えることで悪名高い。
高価な人間の遠隔操作なしにこの混乱を解決するために、VLAモデルは既知のサブスキルを、オリジナルのデモンストレーションを超えてより広範な実行可能行動の集合に構成できるようにする。我々は、このモデルが潜むタスク表現を利用して、ポリシートレーニングのために既存のタスクから直接、新しい物理的に有効なデモを合成するオフラインデータ拡張フレームワークACT-VLA(Action Composal Training for VLA Models)を提案する。
そこで本手法では,手作業によるデータ収集を不要にすることで,トレーニング分布を自動的に拡張し,過度な適合を緩和する。
シミュレーションにおける操作課題に対する我々のアプローチを評価する。
実験により, ベースラインVLAモデルは, もともとの分布オーバーフィッティングによる一般化が不十分であるのに対して, 合成データを用いて訓練したポリシは, 極めて高い成功率を実現し, 既存のタスクを自動実演合成に活用することにより, VLAの一般化を拡大するための効果的でスケーラブルでデータ効率のよい経路を提供することを示した。
関連論文リスト
- Decoupling the Declarative from the Procedural in Vision-Language-Action Models [65.22639791239932]
オブジェクト固有のデモから振る舞いをクローンするように訓練されたポリシーは、そのオブジェクトを超えて一般化されなければならない。
情報フローを再構成した新しいビジョン・ランゲージ・アクションモデルであるw$2$VLAを提案する。
最先端のVLAとは異なり、我々のモジュラーアプローチは知識表現の分離に成功している。
論文 参考訳(メタデータ) (2026-06-19T14:43:58Z) - Reshaping Action Error Distributions for Reliable Vision-Language-Action Models [69.38615670891038]
ロボット操作において、視覚言語アクション(VLA)モデルは、一般化可能でスケーラブルなロボットポリシーを学ぶための有望なパラダイムとして登場した。
連続動作型VLAモデルに焦点をあて、トレーニング中の動作誤差分布を再構成することにより、従来のMSEベースの回帰を超越する。
複数の代表的VLAアーキテクチャ上で、標準、少数ショット、ノイズの多い設定にまたがるアプローチを評価します。
論文 参考訳(メタデータ) (2026-02-04T05:37:09Z) - Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach [78.4812458793128]
動作チャンクの高忠実度検証に軽量な擬数推定器を適用したテスト時間スケーリングフレームワークである textbfTACO を提案する。
我々の手法は、オフライン強化学習(RL)における古典的な反探索原理に似ており、勾配のないため、計算上の大きな恩恵をもたらす。
論文 参考訳(メタデータ) (2025-12-02T14:42:54Z) - Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance [63.33213516925946]
textbfAlign-Then-stEer(textttATE)は,新しいデータ効率,プラグアンドプレイ適応フレームワークである。
我々の研究は、新しいロボットプラットフォームやタスクにVLAモデルをデプロイする実用性を大幅に向上させる、汎用的で軽量なソリューションを提供する。
論文 参考訳(メタデータ) (2025-09-02T07:51:59Z) - Forewarned is Forearmed: Leveraging LLMs for Data Synthesis through Failure-Inducing Exploration [90.41908331897639]
大規模言語モデル(LLM)は、多種多様な高品質なタスク特化データのトレーニングの恩恵を受けている。
本稿では,効果的なトレーニングサンプルを自動生成する新しい手法であるReverseGenを提案する。
論文 参考訳(メタデータ) (2024-10-22T06:43:28Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。