論文の概要: AutodidactWAM: Cross-Modal Self-Distillation from Generated Video to Robot Actions
- arxiv url: http://arxiv.org/abs/2610.08119v1
- Date: Tue, 06 Oct 2026 10:39:42 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-08 02:58:29.940139
- Title: AutodidactWAM: Cross-Modal Self-Distillation from Generated Video to Robot Actions
- Title(参考訳): AutodidactWAM: 生成したビデオからロボット行動へのクロスモーダル自己蒸留
- Abstract要約: 世界行動モデル(WAM)は、観察と指示から将来のビデオとロボットのアクションを共同で生成する。
クローズドループ実ロボット試験を3段階の累積段階(プレグラス,グリップ,ピック・アンド・プレイス)で評価した。
本稿では,遠隔操作を伴わない手動推定器AutodidactWAMを提案する。
- 参考スコア(独自算出の注目度): 2.2935396753701065
- License: http://creativecommons.org/licenses/by-nc-nd/4.0/
- Abstract: World-action models (WAMs) such as Cosmos 3 jointly generate future video and robot actions from an observation and instruction. Adapting one such model with a lightweight LoRA fine-tune to a previously unseen robot, a Unitree G1 humanoid with five-fingered BrainCo hands, exposes a video-action asymmetry: the video renders plausible task executions, while the co-generated action is systematically mis-targeted. We evaluate closed-loop real-robot trials at three cumulative stages: pre-grasp, grasp, and pick-and-place. The native action succeeds only approximately 17%, 10%, and 7% of the time, respectively, and performs worse on held-out objects. We propose AutodidactWAM, a hand-pose estimator trained without teleoperation, followed by inverse kinematics, that runs on the model's generated video to recover action estimates. Paired with the native prediction, these recovered actions provide preferred targets for fine-tuning only the action-related layers, while the generated video is teacher-forced. We compare supervised relabeling with a rectified-flow adaptation of Diffusion-DPO. After one-time embodiment adaptation, self-distillation requires no additional task-specific teleoperation. The recovered-action gate reaches approximately 75%, 47%, and 42% pre-grasp, grasp, and pick-and-place success, compared with 17%, 10%, and 7% for the native action. A hybrid objective combining preference supervision, supervised target fitting, and Cartesian trajectory anchoring (DPO+SFT+DTW) performs best: on Oreo, the training object, it reaches 90% pre-grasp and 20% full-task success; on a held-out object, it reaches 80% and 30%. Plain Flow-DPO reaches 0% success despite 1.000 validation preference accuracy, indicating that the combination of training objectives, rather than the contrastive objective alone, drives the observed gains.
- Abstract(参考訳): コスモス3のような世界アクションモデル(WAM)は、観察と指示から将来のビデオとロボットアクションを共同で生成する。
軽量のLoRAファインチューンと5本指のBrainCoハンドを持つユニツリーG1ヒューマノイドに、そのようなモデルを適用すると、ビデオアクションの非対称性が露呈する。
クローズドループ実ロボット試験を3段階の累積段階(プレグラス,グリップ,ピック・アンド・プレイス)で評価した。
ネイティブアクションは、それぞれ約17%、10%、7%の時間しか成功せず、ホールドアウトされたオブジェクトでさらに悪化する。
本稿では,遠隔操作なしでトレーニングした手動推定器AutodidactWAMを提案する。
ネイティブな予測と合わせて、これらの回復されたアクションはアクション関連レイヤのみを微調整するための望ましいターゲットを提供する一方で、生成されたビデオは教師が強制する。
教師付きレザベリングと拡散DPOの正流適応を比較した。
一度の実施態様の適応の後、自己蒸留は追加のタスク固有の遠隔操作を必要としない。
回収された行動ゲートは約75%, 47%, 42%, 42%, 17%, 10%, 7%であった。
DPO+SFT+DTW(Cartesian trajectory anchoring)とDPO+SFT+DTW(Cartesian trajectory anchoring)を組み合わせたハイブリッドな目的は、Oreoでは、トレーニング対象が90%、フルタスク成功が20%、保留対象が80%と30%である。
通常のフロー-DPOは1.000の検証選好精度にもかかわらず0%の精度で成功し、対照的な目的のみでなくトレーニング目標の組み合わせが観察された利得を駆動していることを示している。
関連論文リスト
- OpenWAM: An Open Framework for Composable World-Action Models [30.981740473743418]
OPENWAMは,共通の因果的ロボット・ビデオ基盤を基盤として構築されたオープンワールドアクションモデリングフレームワークである。
我々は、1万時間以上のビデオで因果的ロボットビデオの事前トレーニングを行い、その後、共有Mixture-of-Transformersアーキテクチャを通じてアクションエキスパートを統合する。
OPENWAMは4つのLIBEROスイートと実世界のバイマニュアルタスクで高い成功率を達成する。
論文 参考訳(メタデータ) (2026-10-06T08:02:12Z) - ACG-WAM: World-Action Modeling via Action-Conditioned Geometric Latent Prediction [4.104352271917983]
アクション・コンディション付き幾何学Joint-Embedding Predictive Architecture (ACG-WAM)
ACG-WAMは、現在の観測と干渉行動からいくつかの地平線における幾何学的特徴を予測する。
50のRoboTwin 2.0タスクでは、ACG-WAMはクリーンなシーンで93.46%の成功を達成している。
論文 参考訳(メタデータ) (2026-10-03T17:52:59Z) - World Action Modeling with Progressive Visual Planning [42.47105321176084]
本稿では,協調して行動を予測するプログレッシブ・ワールド・アクション・モデルであるProWAMについて述べる。
この設計は、大規模アクションフリービデオからサブゴール予測を学ぶことができるため、自然にスケールする。
効率的なアクション生成のために、ProWAMは単一のビデオバックボーンフォワードパスを実行し、スパースサブゴール機能をキャッシュする。
論文 参考訳(メタデータ) (2026-10-01T21:32:29Z) - EVO-WAM: Evolving World Action Models through Video-Action Verification [51.68210232964847]
EVO-WAMは、自身の生成したビデオアクション軌跡から学習することで、WAMを見えないタスクに適応するフレームワークである。
コスモス3の平均成功率は20.0%から76.7%に改善され、56.7%となった。
論文 参考訳(メタデータ) (2026-09-29T17:25:35Z) - GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation [63.853929983870955]
本稿では,ジェニー・エンスペクタ法 2.0 (GE-Act 2.0) について紹介する。
制御指向オートエンコーダ(CoAE)、ワンステップビジュアルプランナー(SVP)、逆ダイナミクスモデル(IDM)を組み合わせている。
それらのコンポーネントは、知識に整合した選択最適化(KASO)で共同で訓練され、記録された行動に適合していると判断された予測された未来のみを選択することで、ミスマッチした監視を減らす。
論文 参考訳(メタデータ) (2026-09-04T17:16:48Z) - Physical Autoregressive Model for Robotic Manipulation without Action Pretraining [65.8971623698511]
我々は、自己回帰ビデオ生成モデルを構築し、物理自己回帰モデル(PAR)を提案する。
PARは、アクション事前トレーニングを必要とせず、物理力学を理解するために、ビデオ事前トレーニングに埋め込まれた世界の知識を活用する。
ManiSkillベンチマークの実験は、PARがPushCubeタスクで100%の成功率を達成したことを示している。
論文 参考訳(メタデータ) (2025-08-13T13:54:51Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。