論文の概要: MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning
- arxiv url: http://arxiv.org/abs/2607.14252v1
- Date: Wed, 15 Jul 2026 18:12:28 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-17 17:01:32.882901
- Title: MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning
- Title(参考訳): MEMORA: 推論と計画のためのエゴセントリックビデオからの身体的アクション記憶
- Authors: Zihao Yu, Xiu Yuan, Chongjie Zhang,
- Abstract要約: ロングホライズンロボットの計画には、将来の目標を解釈可能な具体的体験の記憶が必要である。
私たちは、Embodied Action Memory(EAM)を、後の決定のために永続的なメモリ状態として、そのエクスペリエンスを形成、維持、使用する能力として定式化します。
MEMORAは, 生成・統合・検索ライフサイクルと4つの型付きストアでEMAを実現する。
- 参考スコア(独自算出の注目度): 27.103249755262212
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Long-horizon robot planning requires more than predicting what actions will do next; it also requires memory of the embodied experience that makes future goals interpretable. People do not plan from the present scene alone: they draw on remembered places, object-state changes, prior procedures, and regularities revealed through repeated action. We formulate Embodied Action Memory (EAM) as the capability to form, maintain, and use such experience as a persistent memory state for later decisions. MEMORA realizes EAM with a formation-consolidation-retrieval lifecycle and four typed stores: Environment Memory, Entity Memory, Activity Memory, and Inferred Knowledge. Online editing maintains object identities and state histories as new observations arrive; offline consolidation abstracts repeated experience into reusable procedures and participant-specific regularities. MEMORA-Bench evaluates this lifecycle on 45 hours of EPIC-KITCHENS-100 extension video across 18 participants through memory-grounded planning, including previously unseen goals, and a complementary memory-assessment task. Across four open-weight language models, full MEMORA--combining editing, typed stores, and consolidation--achieves the strongest aggregate results among the evaluated memory conditions. It improves memory-assessment accuracy by up to 20.5 points over the strongest controlled baseline and improves out-of-distribution Robot-Grounded Plan score by up to 16.6% relative. A qualitative two-task robot deployment study further illustrates how memory-grounded language plans can interface with downstream control, while the overall results show that editable, consolidated memory can supply remembered context for robot planning. Project page: https://yuzihaowashu.github.io/MEMORA/
- Abstract(参考訳): ロングホライズンロボットの計画には、次に何をするかを予測する以上のことが必要であり、また将来の目標を解釈可能な具体的体験の記憶も必要である。
人々は、記憶された場所、オブジェクト状態の変更、事前の手順、繰り返しのアクションによって明らかにされる規則など、現在のシーンからのみ計画を立案しません。
EAM(Embodied Action Memory)を、後の決定のために永続的なメモリ状態として生成、維持、使用する能力として定式化します。
MEMORAは、環境記憶、エンティティ記憶、アクティビティ記憶、推論知識の4つの型付きストアで、EAMを実現している。
オフラインの統合は、繰り返しの経験を再利用可能な手順と参加者固有の規則に抽象化する。
MEMORA-Benchはこのライフサイクルを18人の参加者を対象としたEPIC-KITCHENS-100拡張ビデオの45時間で評価している。
4つのオープンウェイト言語モデル、完全なMEMORA編集、型付きストア、コンソリデーションは、評価されたメモリ条件の中で最強の集計結果を取得する。
最強の制御ベースライン上で最大20.5ポイントのメモリ評価精度を改善し、アウト・オブ・ディストリビューション・ロボット・グラウンド・プランのスコアを最大16.6%改善する。
定性的な2タスクロボット配置研究は、メモリグランド言語プランがダウンストリーム制御とどのように相互作用するかをさらに示し、全体的な結果は、編集可能で統合されたメモリが、ロボット計画に記憶されたコンテキストを提供することを示している。
プロジェクトページ:https://yuzihaowashu.github.io/MEMORA/
関連論文リスト
- WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction [72.1620416874118]
マルチモーダルな言語モデルは、長距離エージェントとしてますます多くデプロイされている。
既存のベンチマークは、静的対話上のリコールを測定し、メモリを1つのタスクの精度に分解し、キャプションに対する視覚的な観察を減らす。
マルチモーダルエージェントメモリを,観測可能な4段階ライフサイクルを持つアクションワールドインタラクションループとして定式化する。
論文 参考訳(メタデータ) (2026-05-28T04:27:20Z) - RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies [54.23445842621374]
記憶は、長い水平と歴史に依存したロボット操作にとって重要である。
近年,視覚言語アクション(VLA)モデルにメモリ機構が組み込まれ始めている。
本稿では,VLAモデルの評価と進展のための大規模標準ベンチマークであるRoboMMEを紹介する。
論文 参考訳(メタデータ) (2026-03-04T21:59:32Z) - MEM: Multi-Scale Embodied Memory for Vision Language Action Models [73.3883864595845]
本稿では,マルチスケール・エンボダイドメモリ(MEM)について紹介する。
MEMはビデオベースの短水平メモリをビデオエンコーダで圧縮し、テキストベースの長水平メモリと組み合わせている。
MEMは、キッチンを掃除したり、チーズサンドイッチを焼いたりして、最大15分間のタスクをロボットが実行できるようにする。
論文 参考訳(メタデータ) (2026-03-04T00:03:02Z) - MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation [59.31354761628506]
このようなタスクは本質的にマルコフ的ではないが、主流のVLAモデルはそれを見落としているため、ロボット操作には時間的コンテキストが不可欠である。
本稿では,長距離ロボット操作のためのコグニション・メモリ・アクション・フレームワークであるMemoryVLAを提案する。
本稿では,3つのロボットを対象とした150以上のシミュレーションと実世界のタスクについて評価する。
論文 参考訳(メタデータ) (2025-08-26T17:57:16Z) - RoboMemory: A Brain-inspired Multi-memory Agentic Framework for Interactive Environmental Learning in Physical Embodied Systems [41.89907261427986]
エージェントは、部分的可観測性、空間的推論の制限、高速なマルチメモリ統合など、現実世界の環境において永続的な課題に直面している。
本稿では, 空間, 時間, エピソディック, セマンティックメモリを並列化して, 効率的な長期計画と対話型環境学習を実現する, 脳にインスパイアされたフレームワークであるRoboMemoryを紹介する。
EmbodiedBenchの実験によると、Qwen2.5-VL-72B-Ins上に構築されたRoboMemoryはベースラインを25%上回り、クローズドソース(SOTA)のGemini-1.5を超えている。
論文 参考訳(メタデータ) (2025-08-02T15:39:42Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。