論文の概要: 4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields
- arxiv url: http://arxiv.org/abs/2608.08023v2
- Date: Wed, 12 Aug 2026 13:47:13 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-13 14:28:20.402685
- Title: 4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields
- Title(参考訳): 4D-WAM:軌道場による世界行動モデルへの時空間認識の注入
- Authors: Lishan Yang, Wenxuan Song, Xi Wang, Pingyue Sheng, Zheng Fang, Ziyang Zhou, Junjie He, Haodong Yan, Jiayi Chen, Nan Sun, Qiao Sun, Pengwei Wang, Lingqiao Liu, Yan Wang, Yuxiang Gao, Feras Dayoub, Haoang Li,
- Abstract要約: World Action Models (WAM) は、ビデオ予測とアクション生成を共同でモデル化する。
WAMは2Dピクセル空間でビデオを表現するのが一般的で、ロボットアクションが実行される3D空間との表現ギャップを生じる。
最近の3Dアプローチでは3D情報を導入しているが、3D構造のダイナミクスを完全に活用できない。
本研究では,WAMに3次元軌道場から知識を注入するモデルに依存しないトレーニング戦略である4D-WAMを提案する。
- 参考スコア(独自算出の注目度): 35.51443781145251
- License: http://creativecommons.org/licenses/by-sa/4.0/
- Abstract: Building on recent advances in world models, World Action Models (WAMs) jointly model video prediction and action generation. However, they typically represent videos in 2D pixel space, creating a representation gap with 3D space in which robotic actions are executed. Recent 3D approaches introduce 3D information, but fail to fully exploit the dynamics of 3D structures. In this work, we propose 4D-WAM, a model-agnostic training strategy that injects spatiotemporal knowledge from 3D trajectory fields into WAMs through representation alignment. To this end, we introduce two complementary objectives: 1) motion alignment, which aligns temporal feature variations across adjacent frames and encourages the model to build local 4D awareness during training, and 2) destination alignment, which guides the model to infer the final destination from the source frame by minimizing the gap between their attention-like similarity distributions. Together, these objectives provide both local motion supervision and long-horizon goal guidance, enabling WAMs to learn trajectory-level spatiotemporal representations. Extensive in-distribution and out-of-distribution experiments across different base models demonstrate the model's improvements in spatial understanding, execution precision, robustness, generalization, and versatility.
- Abstract(参考訳): World Action Models (WAMs)は、映像予測とアクション生成を共同でモデル化する。
しかし、通常は2Dピクセル空間でビデオを表現し、ロボットアクションが実行される3D空間との表現ギャップを生じる。
最近の3Dアプローチでは3D情報を導入しているが、3D構造のダイナミクスを完全に活用できない。
本研究では,3次元軌道場からWAMに時空間的知識を注入するモデルに依存しない学習手法である4D-WAMを提案する。
この目的のために,2つの補完的目的を導入する。
1) 運動アライメント(運動アライメント)は、隣接フレーム間の時間的特徴変化を調整し、トレーニング中に局所的な4D認識を構築するようモデルに促す。
2) 宛先アライメントは,アライメントのような類似度分布間のギャップを最小化することにより,モデルにソースフレームから最終目的地を推定させる。
これらの目的は、局所的な運動監督と長期的目標誘導の両方を提供し、WAMが軌跡レベルの時空間表現を学習できるようにする。
異なるベースモデルにまたがる広範囲な流通実験とアウト・オブ・ディストリビューション実験は、空間的理解、実行精度、堅牢性、一般化、汎用性におけるモデルの改善を実証している。
関連論文リスト
- Lift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware Manipulation [76.16065370700488]
We introduced Lift3D-VLA, a unified VLA framework that equips models with explicit 3D point cloud reasoning and possible temporally coherent action generation。
我々は,Lift3D-VLAがMetaWorldとRLBenchの平均成功率を10.8%,11.1%向上したことを示す。
論文 参考訳(メタデータ) (2026-07-07T17:59:47Z) - StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation [6.0744834626758495]
StemVLAは、未来の3D空間知識と歴史的4D表現の両方をアクション予測に明示的に組み込む新しいフレームワークである。
我々は,CALVIN ABC-D ベンチマーク [46] において,StemVLA はタスクの長期化と最先端性能を著しく向上し,XXX の平均シーケンス長を達成できることを示した。
論文 参考訳(メタデータ) (2026-02-27T06:43:37Z) - RAYNOVA: Scale-Temporal Autoregressive World Modeling in Ray Space [51.441415833480505]
RAYNOVAは、二重因果自己回帰フレームワークを使用するシナリオを駆動するための多視点世界モデルである。
相対的なシャーカー線位置符号化に基づいて、ビュー、フレーム、スケールにまたがる等方的時間的表現を構築する。
論文 参考訳(メタデータ) (2026-02-24T08:41:40Z) - UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework [54.337290937468175]
統合された枠組み内での2次元映像と3次元映像の協調モデリングのための自己回帰モデルUniMoを提案する。
本手法は,正確なモーションキャプチャを行いながら,対応する映像と動きを同時に生成することを示す。
論文 参考訳(メタデータ) (2025-12-03T16:03:18Z) - Motion4D: Learning 3D-Consistent Motion and Semantics for 4D Scene Understanding [54.859943475818234]
基礎モデルからの2次元先行を統一された4次元ガウススプラッティング表現に統合する新しいフレームワークであるMotion4Dを提案する。
1) 局所的な一貫性を維持するために連続的に動き場と意味体を更新する逐次最適化,2) 長期的コヒーレンスのために全ての属性を共同で洗練するグローバル最適化,である。
提案手法は,ポイントベーストラッキング,ビデオオブジェクトセグメンテーション,新しいビュー合成など,多様なシーン理解タスクにおいて,2次元基礎モデルと既存の3Dベースアプローチの両方に優れる。
論文 参考訳(メタデータ) (2025-12-03T09:32:56Z) - VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation [54.81449795163812]
時間的コヒーレントなロボット操作のための4次元認識型汎用VLAモデルを開発した。
視覚的特徴を抽出し, 4次元埋め込みのための3次元位置への1次元時間埋め込みを行い, クロスアテンション機構による統一視覚表現に融合する。
この枠組みの中で、デザインされた視覚アクションは、空間的に滑らかで時間的に一貫したロボット操作を共同で行う。
論文 参考訳(メタデータ) (2025-11-21T12:26:30Z) - 3D-VLA: A 3D Vision-Language-Action Generative World Model [68.0388311799959]
最近の視覚言語アクション(VLA)モデルは2D入力に依存しており、3D物理世界の広い領域との統合は欠如している。
本稿では,3次元知覚,推論,行動をシームレスにリンクする新しい基礎モデルのファウンデーションモデルを導入することにより,3D-VLAを提案する。
本実験により,3D-VLAは実環境における推論,マルチモーダル生成,計画能力を大幅に向上することが示された。
論文 参考訳(メタデータ) (2024-03-14T17:58:41Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。