論文の概要: AffordanceWAM: Affordance-Aware Joint World-Action Modeling for Robot Manipulation
- arxiv url: http://arxiv.org/abs/2609.22332v2
- Date: Tue, 22 Sep 2026 05:55:53 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-23 18:04:03.866338
- Title: AffordanceWAM: Affordance-Aware Joint World-Action Modeling for Robot Manipulation
- Title(参考訳): AffordanceWAM:Affordance-Aware Joint World-Action Modeling for Robot Manipulation
- Abstract要約: 汎用的なロボット操作は、シーンがどのように進化するかを予測する必要がある。
アクションラベル付きロボットビデオは、直接制御を監督するが、コストが高く、多様性に制限がある。
本稿では,Scalr Affordance と Affordance Heatmap によるオブジェクトの空き時間を表す,空き時間を考慮した World Action Model を提案する。
- 参考スコア(独自算出の注目度): 61.436480565887905
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Generalizable robot manipulation requires predicting how a scene will evolve, identifying where interactions are feasible, and determining how to act. Action-labeled robot videos directly supervise control but are costly and limited in diversity, whereas egocentric human videos capture diverse interactions but lack robot actions and differ in embodiment and appearance. We introduce AffordanceWAM, an affordance-aware generative World Action Model that represents object-centric spatiotemporal affordance through Scalar Affordance and Affordance Heatmap, within the generated future World. This representation grounds visual prediction in task-relevant objects and interaction regions for action generation, and provides shared interaction targets across human and robot videos. Built on a pretrained video diffusion Transformer, AffordanceWAM uses separately parameterized World and Action Experts, coupled through Masked Joint Self-Attention, to jointly predict future RGB observations, Scalar Affordance fields, Affordance Heatmaps, and continuous robot actions under a unified flow-matching objective. Human videos supervise all three future-World streams, whereas robot trajectories additionally provide action supervision, enabling transfer without human action labels or retargeting. Experiments on RoboCasa, CALVIN ABC$\rightarrow$D, and real-world manipulation demonstrate consistent gains over RGB-only and robot-data-only baselines. Under fixed robot supervision, RoboCasa performance improves monotonically as affordance-annotated human video scales. These results support affordance as an effective interface for both vision-language-action learning and human-to-robot transfer.
- Abstract(参考訳): 一般的なロボット操作には、シーンがどのように進化するかを予測し、相互作用がどこにあるかを特定し、どのように振る舞うかを決定する必要がある。
アクションラベル付きロボットビデオは直接制御を監督するが、コストが高く、多様性に制限がある。
本稿では,Scalr Affordance と Affordance Heatmap によるオブジェクト中心の時空間アベイランスを表す,アベイランス対応のワールドアクションモデル AffordanceWAM を紹介する。
この表現は、タスク関連オブジェクトとアクション生成のためのインタラクション領域の視覚的予測を基盤とし、人間とロボットのビデオ間で共有されたインタラクションターゲットを提供する。
AffordanceWAMは、事前訓練されたビデオ拡散トランスフォーマー上に構築され、Masked Joint Self-Attentionを介して、個別にパラメータ化されたWorldとAction Expertsを使用し、将来のRGB観測、Scalar Affordance Field、Affordance Heatmaps、およびフローマッチングの目的の下での連続ロボットアクションを共同で予測する。
人間のビデオは、未来の3つのストリームを監督するが、ロボットの軌跡は、人間のアクションラベルやリターゲティングを使わずに、アクションを監督する。
RoboCasa、CALVIN ABC$\rightarrow$D、実世界の操作実験は、RGBのみのベースラインとロボットデータのみのベースラインよりも一貫した利得を示している。
固定されたロボットの監督の下では、RoboCasaのパフォーマンスは、割当アノテートされた人間のビデオスケールとして単調に改善される。
これらの結果は、視覚-言語-行動学習と人間-ロボット間の移動の両方に有効なインタフェースとして、手頃さを支持する。
関連論文リスト
- AnyWorld: Factorized Egocentric World Models for Cross-Embodiment Generalization [88.25615187835321]
我々は,クロス・エボディメント・ワールド・モデリング・フレームワークであるAnyWorldを提案する。
私たちのモデルは、アクション、カメラ、エンボディメントへのインタラクションを分解します。
大規模なヒューマンインタラクション事前トレーニングでモデルをトレーニングする。
論文 参考訳(メタデータ) (2026-08-29T12:55:15Z) - H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models [56.60098789696351]
H2R-Benchは、人間とロボットの相互操作ビデオ生成を評価するためのベンチマークである。
各ベンチマークインスタンスには、人間のデモビデオ、ターゲットの実施制約、ソース地上アノテーションが含まれている。
我々は、6つの操作ファミリーと2つのロボットエボディメントで11の最先端のビデオ生成モデルをベンチマークした。
論文 参考訳(メタデータ) (2026-08-13T10:14:33Z) - Robot-Factored World Models via Robot Rendering [25.12876640584098]
アクション条件付きビデオワールドモデルは、初期観測と行動信号から将来の観測を予測する。
本研究では,ロボット特異的な2つの要素を世界モデル外へ移動させるロボット駆動世界モデルを提案する。
実験により, レンダリングインタフェースはベクトル条件付きベースラインより優れており, 未知のロボットエンボディメント推論に一般化されていることがわかった。
論文 参考訳(メタデータ) (2026-07-24T17:59:57Z) - WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos [38.959415159248046]
WALAは、アクションラベル付きデモとアクションフリービデオの両方から実行可能な潜在アクションを学ぶためのフレームワークである。
WALAはRoboTwin上で高いパフォーマンスを実現し,RoboCasa上での最先端の成果を平均75.2%で実現している。
論文 参考訳(メタデータ) (2026-07-13T11:02:04Z) - MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training [102.850162490626]
人間のロボットによる相互模倣事前学習による視覚-言語-行動モデルであるMiVLAを提案する。
MiVLAは、最先端のVLAよりも優れた、強力な改良された一般化能力を実現する。
論文 参考訳(メタデータ) (2025-12-17T12:59:41Z) - MimicDreamer: Aligning Human and Robot Demonstrations for Scalable VLA Training [40.45924128424013]
低コストな人間によるデモンストレーションをロボットで使用可能な監視に変換するフレームワークであるMimicDreamerを提案する。
視覚的アライメントのために,高忠実度ロボットデモビデオを生成するビデオ拡散モデルH2R Alignerを提案する。
視点安定化のためにEgoStabilizerを提案する。
動作アライメントのために,人間の手の動きをロボットフレームにマッピングし,制約付き逆運動学解法を適用する。
論文 参考訳(メタデータ) (2025-09-26T11:05:10Z) - Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation [65.46610405509338]
我々は、ゼロショットロボット操作を可能にする汎用的な目標条件ポリシーを学習することを目指している。
私たちのフレームワークであるTrack2Actは、ゴールに基づいて将来のタイムステップで画像内のポイントがどのように動くかを予測する。
学習したトラック予測を残留ポリシーと組み合わせることで,多種多様な汎用ロボット操作が可能となることを示す。
論文 参考訳(メタデータ) (2024-05-02T17:56:55Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。