論文の概要: UniWAM: Unified World-Action Model
- arxiv url: http://arxiv.org/abs/2610.02054v2
- Date: Sat, 03 Oct 2026 05:46:02 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-07 04:43:28.525964
- Title: UniWAM: Unified World-Action Model
- Title(参考訳): UniWAM:Unified World-Action Model
- Abstract要約: UniWAMは物理推論器、世界生成器、行動予測器を統合し、物理世界、視覚生成、行動予測のセマンティック理解を共同で学習する。
UniWAMは、分散性能、一般化、命令追従、長期タスク実行など、複数の評価において最先端(SOTA)性能を達成する。
- 参考スコア(独自算出の注目度): 18.452844176630276
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semantic understanding and reasoning under distribution shifts. We introduce UniWAM, a unified architecture that integrates a physical reasoner, a world generator, and an action predictor to jointly learn semantic understanding of the physical world, visual generation, and action prediction. To ensure the quality of the training data, we developed a rigorous data cleaning and annotation pipeline for both human egocentric data and robot data. To adapt the vision-language component to embodied tasks while preserving its inherited language capabilities, we represent low-level actions in natural language and introduce a pre-training recipe that assigns complementary supervision from visual question answering (VQA) data, human egocentric data, and robot demonstrations to the appropriate model components. During post-training, future visual noise augmentation reduces reliance on precise future predictions, while history-conditioned flow matching uses encoded action history to initialize action generation. Together, these designs significantly reduce denoising steps while maintaining performance. UniWAM achieves state-of-the-art (SOTA) performance across multiple evaluations, including in-distribution performance, robustness, generalization, instruction following, and long-horizon task execution. Furthermore, we uncover a log-linear scaling law of unified human-robot co-training, demonstrating the effectiveness of large-scale pre-training on a mixture of human and robot data.
- Abstract(参考訳): 視覚言語アクションモデルは、事前訓練された視覚言語モデルの理解と推論能力の恩恵を受けるが、アクションのみの監督は、世界力学の限られた基盤を提供する。
逆に、世界行動モデルはビデオ生成モデルから時空間的先行を継承するが、分配シフトの下では意味的理解や推論に制限される。
我々は、物理世界、視覚生成、行動予測のセマンティック理解を共同で学習するための物理推論器、世界生成器、行動予測器を統合した統一アーキテクチャUniWAMを紹介する。
トレーニングデータの質を確保するため,人間中心のデータとロボットデータの両方を対象とした厳密なデータクリーニングとアノテーションパイプラインを開発した。
視覚言語コンポーネントをタスクの具現化に適応させるため,自然言語の低レベル動作を表現し,視覚的質問応答(VQA)データ,人間中心データ,ロボットのデモを適切なモデルコンポーネントに割り当てる事前学習レシピを導入する。
後のトレーニングでは、将来の視覚ノイズの増大は正確な将来の予測への依存を減らし、履歴条件付きフローマッチングは、エンコードされたアクション履歴を使用してアクション生成を初期化する。
これらの設計により、性能を維持しながらデノイングのステップが大幅に削減される。
UniWAMは、分散性能、堅牢性、一般化、命令追従、長期タスク実行など、複数の評価で最先端(SOTA)性能を実現している。
さらに,人間とロボットの混在による大規模事前学習の有効性を実証し,統合されたロボット協調学習の対数線スケーリング法則を明らかにする。
関連論文リスト
- EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action [65.7650775259592]
視覚言語行動(VLA)ポリシーは意味理解を重視し、世界行動モデル(WAM)は環境力学の予測表現を学習する。
本研究では,アクション中心の統一型エンボディモデルであるEWAMについて述べる。
論文 参考訳(メタデータ) (2026-09-30T15:35:53Z) - mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs [5.109732854501585]
そこで我々は,事前学習したインターネットスケールのビデオモデルと,その潜在表現に条件付けされたフローマッチングに基づくアクションデコーダを組み合わせた,新しいビデオ・アクション・モデル(VAM)を提案する。
提案手法は,シミュレーションおよび実世界のロボット操作タスクにおける最先端性能を実現し,サンプル効率を10倍,収束速度を2倍向上させる。
論文 参考訳(メタデータ) (2025-12-17T18:47:31Z) - Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations [26.678553477485362]
本稿では,ロボット操作に適応しながら,事前学習した特徴をよりよく保存するフレームワークを提案する。
提案手法では, (i) 事前学習された特徴を保持するために, 凍結したビジョンを持つデュアルエンコーダ設計と, (ii) モデルの事前学習領域に整合した文字列に連続的なアクションを投入する文字列ベースのアクショントークン化器, (iii) ロボットのデモンストレーションと,空間的推論とアプライアンスを強調する視覚言語データセットを組み合わせた協調学習戦略の3つのコンポーネントを導入している。
論文 参考訳(メタデータ) (2025-09-14T20:08:56Z) - DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge [41.030494146004806]
本稿では,逆動力学モデリングを実現するために,包括的世界知識予測を統合した新しいVLAフレームワークであるDreamVLAを提案する。
DreamVLAは、動的領域誘導の世界知識予測を導入し、空間的および意味的な手がかりと統合し、アクション計画のためのコンパクトで包括的な表現を提供する。
実世界とシミュレーション環境での実験では、ドリームVLAが実際のロボットタスクで76.7%の成功率を達成したことが示されている。
論文 参考訳(メタデータ) (2025-07-06T16:14:29Z) - A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning [67.72413262980272]
事前訓練された視覚モデル(PVM)は現代のロボティクスの基本であるが、その最適構成は定かではない。
セマンティック・ボトルネックを導入してオブジェクト中心の表現を誘導する手法であるSlotMIMを開発した。
提案手法は,画像認識,シーン理解,ロボット学習評価において,従来の作業よりも大幅に改善されている。
論文 参考訳(メタデータ) (2025-03-10T06:18:31Z) - Visual Grounding Helps Learn Word Meanings in Low-Data Regimes [47.7950860342515]
現代のニューラル言語モデル(LM)は、人間の文の生成と理解をモデル化するための強力なツールである。
しかし、これらの結果を得るためには、LMは明らかに非人間的な方法で訓練されなければならない。
より自然主義的に訓練されたモデルは、より人間らしい言語学習を示すのか?
本稿では,言語習得における重要なサブタスクである単語学習の文脈において,この問題を考察する。
論文 参考訳(メタデータ) (2023-10-20T03:33:36Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。