論文の概要: DIM-WAM: World-Action Modeling with Diverse Historical Event Memory
- arxiv url: http://arxiv.org/abs/2606.27677v1
- Date: Fri, 26 Jun 2026 03:17:39 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-29 18:24:25.357725
- Title: DIM-WAM: World-Action Modeling with Diverse Historical Event Memory
- Title(参考訳): DIM-WAM: 各種履歴イベントメモリを用いたワールド・アクション・モデリング
- Authors: Kai Wang, Zhaopeng Gu, Yixiang Chen, Yuan Xu, Qisen Ma, Peng Su, Zhaowen Li, Yan Huang, Liang Wang,
- Abstract要約: 世界行動モデルは、将来の視覚状態と行動を共同で予測することで、有望なロボット操作性能を示す。
マルチスケールの歴史的文脈、ローカルな未来のダイナミクス、グローバルなタスク進捗を統合したメモリ拡張ワールドアクションモデルであるDiM-WAMを紹介する。
RMBenchでは、DiM-WAMが平均28.4%、LingBot-VAが69.8%、メモリメモリのMem-0ベースラインが42.0%である。
- 参考スコア(独自算出の注目度): 23.350949759638056
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: World-action models have shown promising robot-manipulation performance by jointly predicting future visual states and actions. However, existing methods mainly rely on short-term history and short-horizon future prediction, which is insufficient for long-horizon tasks whose correct execution depends on earlier observations and task progress. Such temporally dependent tasks require effective use of complementary temporal information, including recent local context, cross-stage historical events, immediate future dynamics, and global task progress. To address long-term forgetting and poor awareness of the global task state, we introduce DiM-WAM, a memory-augmented world-action model that integrates multi-scale historical context, local future dynamics, and global task progress. The memory extracts compact visual event information from real observations, updates multiple memory banks through independent similarity-based merging, and then reads the bank-identity- and time-embedded long-term context to condition video and action denoising. A progress-supervision objective further encourages memory tokens to encode not only completed historical events but also the current task stage and its implications for the remaining task. On RMBench, DiM-WAM raises average success from 28.4% with LingBot-VA to 69.8%, exceeding the explicit-memory Mem-0 baseline at 42.0%. On four real-world Franka tasks, it improves average stage success from 70.7% to 91.5% and full-task success from 52.5% to 80.0%. Project page: https://wangkai-casia.github.io/dim-wam/{\texttt{https://wangkai-casia.github.io/dim-wam/}}.
- Abstract(参考訳): 世界行動モデルは、将来の視覚状態と行動を共同で予測することで、有望なロボット操作性能を示す。
しかし,従来の手法は短期的履歴と短期的将来予測に大きく依存しており,早期の観察や課題進捗に依存する長期的タスクには不十分である。
このような時間依存的なタスクは、最近の局地的な状況、異段階の歴史的出来事、将来のダイナミクス、グローバルなタスク進捗など、補完的な時間的情報を効果的に活用する必要がある。
グローバルタスク状態の長期的忘れと認識の低さに対処するために,マルチスケールな歴史的文脈,ローカルな未来のダイナミクス,グローバルなタスク進捗を統合したメモリ拡張ワールドアクションモデルであるDiM-WAMを導入する。
メモリは、実際の観測からコンパクトな視覚イベント情報を抽出し、独立した類似性ベースのマージを通じて複数のメモリバンクを更新し、バンクアイデンティティとタイムエンベッドされた長期コンテキストを読み出して、条件ビデオとアクションデノーミングを行う。
プログレス・スーパービジョンの目的は、メモリトークンが、完了した履歴イベントだけでなく、現在のタスクステージとその残りのタスクへの含意もエンコードすることを奨励する。
RMBenchでは、DiM-WAMが平均28.4%、LingBot-VAが69.8%、メモリメモリのMem-0ベースラインが42.0%である。
4つの実世界のフランカタスクでは、平均的なステージ成功率を70.7%から91.5%に、フルタスクの成功率を52.5%から80.0%に改善する。
プロジェクトページ:https://wangkai-casia.github.io/dim-wam/{\texttt{https://wangkai-casia.github.io/dim-wam/}}。
関連論文リスト
- KEMO: Event-Driven Keyframe Memory for Long-Horizon Robot Manipulation with VLA Policies [14.253600207124194]
メモリはこの曖昧さに対処し、ポリシーが実行履歴からタスクの進捗を推測できるようにする。
本稿では,VLA ポリシーのタスク関連状態変化に関連するキネマティクスを自動的に保存する,軽量なプラグインメモリフレームワーク KEMO を提案する。
830ステップから2846ステップ(28秒から95秒)の軌道長と2~6サブタスクにまたがる実世界のマルチアーム操作タスクにおけるKEMOの評価を行った。
論文 参考訳(メタデータ) (2026-06-22T16:57:43Z) - MemoryWAM: Efficient World Action Modeling with Persistent Memory [80.90899269062128]
本稿では,効率的な永続メモリを持つ世界アクションモデルであるMemoryWAMを紹介する。
調整された注意機構は、詳細な短期コンテキストと圧縮された長期コンテキストの両方の検索を可能にする。
MemoryWAMは、シミュレーションと実世界の両方において、強力な視覚言語アクション(VLA)とWAMベースラインを上回っている。
論文 参考訳(メタデータ) (2026-06-18T17:59:51Z) - Mem-World: Memory-Augmented Action-Conditioned World Models for Persistent Robot Manipulation [55.42006264038458]
アクション条件付き世界モデルは、ロボット学習の有望なパラダイムとして登場した。
Mem-Worldはメモリ拡張されたマルチビューアクション条件の世界モデルである。
W-VMemは4次元手首ビュー中心のサーベイルインデクシングメモリで、歴史的観測を時間的に変化する表面要素に固定する。
論文 参考訳(メタデータ) (2026-06-17T11:42:00Z) - HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy [61.668591984635846]
HAMLETは、行動予測中の歴史的状況に対応するためにビジョン・ランゲージ・アクションモデルを適用するためのフレームワークである。
HAMLETは、最先端のVLAを履歴認識ポリシーに変換することに成功していることを示す。
GR00T N1.5に加えて、HAMLETは歴史に依存した実世界のタスクで平均76.4%の成功率を達成した。
論文 参考訳(メタデータ) (2025-10-01T09:15:52Z) - FindingDory: A Benchmark to Evaluate Memory in Embodied Agents [49.18498389833308]
本研究では,Habitatシミュレータに長距離エンボディタスクのための新しいベンチマークを導入する。
このベンチマークは、持続的なエンゲージメントとコンテキスト認識を必要とする60タスクにわたるメモリベースの機能を評価する。
論文 参考訳(メタデータ) (2025-06-18T17:06:28Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。