論文の概要: DynaWM: A Base-VLA-Guided World Foundation Model for Moving-Object Manipulation
- arxiv url: http://arxiv.org/abs/2607.02604v1
- Date: Wed, 01 Jul 2026 13:16:44 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-07 22:26:29.346782
- Title: DynaWM: A Base-VLA-Guided World Foundation Model for Moving-Object Manipulation
- Title(参考訳): DynaWM:移動物体操作のためのベースVLA誘導ワールドファンデーションモデル
- Abstract要約: 移動物体操作のためのベースVLA誘導世界基盤モデルDynaWMを提案する。
微調整されたベースVLAチェックポイントでは、SmolVLA、X-VLA、0、0.5よりも7.19、45.31、1.88、10.94のパーセンテージ改善が達成されている。
アブレーション実験では、視覚的エンコーディングによって成功率が27.50%向上し、アクションコンディショニングが取り除かれた場合、成功率が45.44%低下する。
- 参考スコア(独自算出の注目度): 14.968468495338477
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Although vision-language-action (VLA) models have received widespread attention, many challenges remain in manipulating dynamic moving objects. In most existing approaches, end-to-end forward or inverse dynamics models, i.e., world models, are incorporated into high-performance base VLA architectures, which may degrade the performance of well-pretrained base VLA models due to inappropriate fine-tuning. In this paper, we propose DynaWM, a base-VLA-guided world foundation model that adapts to a wide variety of fine-tuned and coarse-tuned base-VLA checkpoints for moving-object manipulation. DynaWM uses a Mamba-3-based action encoder to encode the base action chunk produced by the base VLA into an action-conditioning representation, a V-JEPA 2.1 vision encoder to extract features from multi-view observation history, and a proprioceptive state encoder to encode robotic-arm proprioceptive states. These feature representations jointly condition a flow-matching DiT to regenerate motion-aware action trajectories for moving-object manipulation. For systematic evaluation, we construct the DynaGrasp-32 benchmark, covering six categories of moving-object manipulation tasks, including velocity variation, trajectory variation, and multi-object manipulation, as well as the DynaGrasp-1600 dataset, which consists of 32 scenarios, 1,600 demonstration trajectories, and approximately 1.53M images. For fine-tuned base-VLA checkpoints, DynaWM achieves percentage improvements of 7.19, 45.31, 1.88, and 10.94 over SmolVLA, X-VLA, π0, and π0.5, respectively. For coarse-tuned base-VLA checkpoints, performance increases by 35.13, 44.06, 35.69, and 26.13 percentage, respectively. Ablation experiments show that visual encoding enhances success by 27.50%, while reducing success by 45.44% if action conditioning is removed.
- Abstract(参考訳): 視覚言語アクション(VLA)モデルは広く注目を集めているが、動的に動く物体を操作する上で多くの課題が残っている。
既存のほとんどのアプローチでは、エンド・ツー・エンド・エンド・フォワードまたはリバース・ダイナミクス・モデル、すなわちワールド・モデルが高性能ベースVLAアーキテクチャに組み込まれており、不適切な微調整のため、よく訓練されたベースVLAモデルの性能が低下する可能性がある。
本稿では,移動物体操作のための多種多様な微調整および粗調整のベースVLAチェックポイントに対応するベースVLA誘導世界基盤モデルDynaWMを提案する。
DynaWMは、Mamba-3ベースのアクションエンコーダを使用して、ベースVLAによって生成されたベースアクションチャンクをアクションコンディショニング表現にエンコードする。
これらの特徴表現は、移動物体操作のための動き認識動作軌跡を再生するフローマッチングDiTを共同条件とする。
系統的な評価のために,DynaGrasp-32ベンチマークを構築し,速度変動,軌道変動,多目的操作を含む移動物体操作の6つのカテゴリと,32のシナリオ,1,600の実証軌道,約1.53万の画像からなるDynaGrasp-1600データセットについて検討した。
微調整されたベースVLAチェックポイントに対して、DynaWMはSmolVLA、X-VLA、π0、π0.5よりも7.19、45.31、1.88、10.94のパーセンテージ改善を達成する。
粗調整のベースVLAチェックポイントでは、それぞれ35.13、44.06、35.69、26.13の割合でパフォーマンスが向上した。
アブレーション実験により、視覚的エンコーディングは成功率を27.50%向上させ、アクション条件が取り除かれた場合、成功率を45.44%低下させることが示された。
関連論文リスト
- GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation [49.16739604572808]
VLA(Vision-Language-Action)モデルは、強力なベンチマークパフォーマンスを実現するが、目に見えないオブジェクトによる現実世界のデプロイに苦労する。
これは、統合幾何認識の操作表現が欠如していることに起因していると我々は主張する。
一般化可能なロボット操作のための統合幾何認識行動表現を学習するためのVLAフレームワークであるGEAR-VLAを提案する。
論文 参考訳(メタデータ) (2026-06-07T09:23:16Z) - ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models [46.57405778313275]
VLAモデル(Vision-Language-Action Model)は、汎用的なロボット制御のための強力なパラダイムである。
ElegantVLAは、モデル内動的計算スケジューリングによってVLAモデルを高速化するプラグイン位相適応推論フレームワークである。
GR00TとCogACTの実験は最大2.55倍と3.77倍のスピードアップを実現し、6つの現実世界のGR00TタスクではElegantVLAは計算を2.18倍に削減し、制御周波数を13.8Hzから26.3Hzに引き上げた。
論文 参考訳(メタデータ) (2026-05-28T06:33:05Z) - RotVLA: Rotational Latent Action for Vision-Language-Action Model [54.22746299071677]
本稿では,連続的な回転潜在動作表現に基づくVLAフレームワークであるRotVLAを紹介する。
潜在作用はSO(n) の元としてモデル化され、連続性、構成性、および実世界の作用力学と整合した構造的幾何学を提供する。
RotVLAはVLMバックボーンとフローマッチングアクションヘッドで構成される。
論文 参考訳(メタデータ) (2026-05-13T11:58:02Z) - DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation [52.83157499300261]
時間的推論と閉ループ適応を統合した動的オブジェクト操作のフレームワークであるDynamicVLAを提案する。
我々は、自動データ収集パイプラインでスクラッチから構築されたDynamic Object Manipulationベンチマークを紹介します。
広範囲な評価は、応答速度、知覚、一般化の顕著な改善を示している。
論文 参考訳(メタデータ) (2026-01-29T18:59:51Z) - TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies [95.30717188630432]
VLAモデルの行動予測のための時空間認識を容易にするために,視覚的トレースプロンプトを導入する。
我々は,これまでに収集した150Kロボット操作トラジェクトリのデータセットに基づいてOpenVLAを微調整し,新しいTraceVLAモデルを開発した。
4B Phi-3-Vision に基づくコンパクトな VLA モデルを提案する。
論文 参考訳(メタデータ) (2024-12-13T18:40:51Z) - CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation [100.25567121604382]
VLA(Vision-Language-Action)モデルは、言語誘導されたタスクの実行と、目に見えないシナリオへの一般化の観点から、ロボット操作を改善した。
VLM(Vision-Language-Models)に基づく新しい高度なVLAアーキテクチャを提案する。
我々のモデルはタスクパフォーマンスにおいて既存のVLAをはるかに上回るだけでなく、新しいロボットへの顕著な適応と、見えないオブジェクトや背景への一般化も示している。
論文 参考訳(メタデータ) (2024-11-29T12:06:03Z) - An Empirical Study of Training End-to-End Vision-and-Language
Transformers [50.23532518166621]
我々はMETER(textbfMultimodal textbfEnd-to-end textbfTransformtextbfER)を提案する。
具体的には、視覚エンコーダ(例えば、CLIP-ViT、Swin変換器)、テキストエンコーダ(例えば、RoBERTa、DeBERTa)、マルチモーダルフュージョン(例えば、マージアテンション対共振器)である。
論文 参考訳(メタデータ) (2021-11-03T17:55:36Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。