論文の概要: ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training
- arxiv url: http://arxiv.org/abs/2609.00188v1
- Date: Mon, 31 Aug 2026 18:09:52 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-02 16:31:35.92825
- Title: ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training
- Title(参考訳): ZimaBlue: スケーラブルなビデオ事前トレーニングを通じて、一般化可能なワールドアクションモデルを進化させる
- Authors: Xionghao Wu, Yijun Yang, Shiyang Zhou, Haoze Sun, Jianhui Liu, Songsong Yu, Jiyao Zhang, Wenbo Li, Bo Wang, Guoqing Ma, Lin Song, Renjie Liao, Shenghe Zheng, Wei Tang, Xiaojuan Qi, Yanwei Li, Yuan Zhang, Zhuotao Tian, Haoyang Huang, Nan Duan,
- Abstract要約: 大規模ビデオから一般化可能なワールドアクションモデルを学ぶためのフレームワークであるZimaBlueを紹介した。
ZimaBlueは3段階のトレーニングカリキュラムに従っており、まず大規模な人間とロボットの自我中心のビデオで因果的エンボディドビデオの事前トレーニングを行う。
ZimaBlueはさらに複数のベンチマークで高いパフォーマンスを実現している。
- 参考スコア(独自算出の注目度): 93.39352617578857
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Robotic manipulation faces a fundamental scaling challenge: robust generalization demands broad physical experience, yet action-labeled robot trajectories are expensive to collect and inherently limited in diversity. Egocentric videos offer a far more scalable source of embodied experience, capturing object interactions, contact dynamics, tool use, and long-horizon behaviors across diverse environments. The central challenge is how to convert this abundant but action-free experience into effective robot control. We introduce ZimaBlue, a scalable framework for learning generalizable World Action Models (WAMs) from large-scale video. ZimaBlue follows a three-stage training curriculum: it first performs causal embodied video pre-training on large-scale human and robot egocentric videos, then grounds the learned visual dynamics in heterogeneous robot trajectories through video-action mid-training with a unified action representation, and finally specializes the model to a target robot for deployment. To make generative WAMs practical for real-time control, ZimaBluefurther adopts an asynchronous Slow-Fast dual-system architecture, where a high-capacity Slow world model provides generalizable spatiotemporal representations and a lightweight Fast branch enables 30 Hz action prediction on NVIDIA RTX 4090. On real-robot zero-shot evaluations, scaling from target-robot data alone to over 120,000 hours of embodied video improves success from 36.1% to 77.8%. ZimaBlue further delivers strong performance across multiple benchmarks, with particularly pronounced gains on unseen tasks.
- Abstract(参考訳): 堅牢な一般化は、幅広い物理的経験を必要とするが、アクションラベル付きロボット軌道は、多様性を収集し、本質的に制限する費用がかかる。
エゴセントリックなビデオは、オブジェクトのインタラクション、コンタクトダイナミクス、ツールの使用、さまざまな環境における長時間の振る舞いをキャプチャする、はるかにスケーラブルなエンボディードエクスペリエンスのソースを提供する。
中心となる課題は、この豊富なアクションフリー体験を効果的なロボット制御に変換する方法だ。
本稿では,大規模ビデオから一般化可能なワールドアクションモデル(WAM)を学習するためのスケーラブルなフレームワークであるZimaBlueを紹介する。
ZimaBlueは3段階のトレーニングカリキュラムに従っており、まず大規模な人間とロボットの自我中心のビデオで因果的エンボディドビデオの事前トレーニングを行い、次に、統一されたアクション表現でビデオアクションの途中のトレーニングを通じて、学習された視覚力学を異種ロボットの軌道に接地し、最終的にモデルをターゲットロボットに展開する。
ZimaBluefurtherは、リアルタイム制御に実用的な生成WAMを実現するために、非同期のSlow-Fastデュアルシステムアーキテクチャを採用し、高容量のSlow Worldモデルが一般化可能な時空間表現を提供し、軽量のFastブランチはNVIDIA RTX 4090上で30Hzの動作予測を可能にする。
リアルロボットゼロショットの評価では、ターゲットロボットのデータのみから12000時間以上のエンボディドビデオへのスケーリングは、36.1%から77.8%に改善する。
ZimaBlueはさらに複数のベンチマークで高いパフォーマンスを実現している。
関連論文リスト
- AnyWorld: Factorized Egocentric World Models for Cross-Embodiment Generalization [88.25615187835321]
我々は,クロス・エボディメント・ワールド・モデリング・フレームワークであるAnyWorldを提案する。
私たちのモデルは、アクション、カメラ、エンボディメントへのインタラクションを分解します。
大規模なヒューマンインタラクション事前トレーニングでモデルをトレーニングする。
論文 参考訳(メタデータ) (2026-08-29T12:55:15Z) - Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining [28.30092786035367]
DeFIはビジュアルフォワードと逆ダイナミクスを分離し、各データソースを利用するための新しいフレームワークである。
今後の予測のために,多種多様な人・ロボットビデオで事前訓練された一般フォワード・ダイナミクス・モデル(GFDM)と,ラベルなしビデオ遷移から潜伏行動を予測するための自己教師付き学習によって訓練された一般逆ダイナミクス・モデル(GIDM)を紹介する。
CALVIN ABC-D と SimplerEnv の実験では、DeFI は CALVIN の平均タスク長 4.51 に達し、SimplerEnv-Frac は 51.2% 成功した。
論文 参考訳(メタデータ) (2026-03-27T17:20:10Z) - Large Video Planner Enables Generalizable Robot Control [117.49024534548319]
汎用ロボットは、様々なタスクや環境にまたがって一般化する意思決定モデルを必要とする。
最近の研究は、マルチモーダル大言語モデル(LM)をアクション出力で拡張し、視覚-アクション(VLA)システムを構築することで、ロボット基盤モデルを構築している。
本稿では,ロボット基礎モデル構築における主要なモダリティとして,大規模ビデオ事前学習を用いるための代替パラダイムについて検討する。
論文 参考訳(メタデータ) (2025-12-17T18:35:54Z) - MimicDreamer: Aligning Human and Robot Demonstrations for Scalable VLA Training [40.45924128424013]
低コストな人間によるデモンストレーションをロボットで使用可能な監視に変換するフレームワークであるMimicDreamerを提案する。
視覚的アライメントのために,高忠実度ロボットデモビデオを生成するビデオ拡散モデルH2R Alignerを提案する。
視点安定化のためにEgoStabilizerを提案する。
動作アライメントのために,人間の手の動きをロボットフレームにマッピングし,制約付き逆運動学解法を適用する。
論文 参考訳(メタデータ) (2025-09-26T11:05:10Z) - VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation [53.63540587160549]
VidBotは、WildのモノクルなRGBのみの人間ビデオから学習した3Dアベイランスを使って、ゼロショットロボット操作を可能にするフレームワークである。
VidBotは、人間の日常的なビデオを利用してロボットの学習をよりスケーラブルにする。
論文 参考訳(メタデータ) (2025-03-10T10:04:58Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。