論文の概要: EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
- arxiv url: http://arxiv.org/abs/2605.09874v1
- Date: Mon, 11 May 2026 01:59:59 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-05-12 23:28:50.467018
- Title: EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
- Title(参考訳): EgoMemReason: 長距離エゴ中心ビデオ理解のためのメモリ駆動推論ベンチマーク
- Authors: Ziyang Wang, Yue Zhang, Shoubin Yu, Ce Zhang, Zengqi Zhao, Jaehong Yoon, Hyunji Lee, Gedas Bertasius, Mohit Bansal,
- Abstract要約: EgoMemReasonは、メモリ駆動推論を通じて、1週間のエゴセントリックなビデオ理解を体系的に評価する。
EgoMemReasonには3つのメモリタイプと6つのコア課題に関する500の質問が含まれている。
EgoMemReasonをMLLMとエージェントフレームワークにまたがる17の手法で評価する。
- 参考スコア(独自算出の注目度): 89.26501160264199
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Next-generation visual assistants, such as smart glasses, embodied agents, and always-on life-logging systems, must reason over an entire day or more of continuous visual experience. In ultra-long video settings, relevant information is sparsely distributed across hours or days, making memory a fundamental challenge: models must accumulate information over time, recall prior states, track temporal order, and abstract recurring patterns. However, existing week-long video benchmarks are primarily designed for perception and recognition, such as moment localization or global summarization, rather than reasoning that requires integrating evidence across multiple days. To address this gap, we introduce EgoMemReason, a comprehensive benchmark that systematically evaluates week-long egocentric video understanding through memory-driven reasoning. EgoMemReason evaluates three complementary memory types: entity memory, tracking how object states evolve and change across days; event memory, recalling and ordering activities separated by hours or days; and behavior memory, abstracting recurring patterns from sparse, repeated observations over the whole week period. EgoMemReason comprises 500 questions across three memory types and six core challenges, with an average of 5.1 video segments of evidence per question and 25.9 hours of memory backtracking. We evaluate EgoMemReason on 17 methods across MLLMs and agentic frameworks, revealing that even the best model achieves only 39.6% overall accuracy. Further analysis shows that the three memory types fail for distinct reasons and that performance degrades as evidence spans longer temporal horizons, revealing that long-horizon memory remains far from solved. We believe EgoMemReason establishes a strong foundation for evaluating and advancing long-context, memory-aware multimodal systems.
- Abstract(参考訳): スマートグラス、エンボディエージェント、常時オンのライフログシステムといった次世代のビジュアルアシスタントは、1日以上にわたって連続的な視覚体験を推論しなければならない。
超長期のビデオ設定では、関連する情報は時間や数日に分散し、メモリが根本的な課題となる。
しかし、既存の1週間のビデオベンチマークは主に、複数の日にわたって証拠を統合する必要のある推論ではなく、モーメントローカライゼーションやグローバルな要約のような認識と認識のために設計されている。
このギャップに対処するために、メモリ駆動推論による1週間のエゴセントリックなビデオ理解を体系的に評価する包括的なベンチマークであるEgoMemReasonを紹介した。
EgoMemReason氏は、エンティティメモリ、オブジェクト状態の進化と変化の追跡、イベントメモリ、数時間または数日で分離されたアクティビティのリコールと順序付け、行動記憶、スパースから繰り返しパターンを抽象化する行動記憶の3つの補完記憶タイプを評価している。
EgoMemReasonには3つのメモリタイプと6つのコア課題に500の質問があり、平均5.1のエビデンスセグメントと25.9時間のメモリバックトラックがある。
EgoMemReasonをMLLMとエージェントフレームワークにまたがる17の手法で評価したところ、最高のモデルでさえ全体的な精度は39.6%に過ぎなかった。
さらなる分析では、3つのメモリタイプが異なる理由で失敗し、エビデンスが時間的地平線にまたがるほど性能が低下していることが示され、長い水平メモリの解決には程遠いことが判明した。
我々は、EgoMemReasonが、長期コンテキスト、メモリ対応マルチモーダルシステムの評価と発展のための強力な基盤を確立していると信じている。
関連論文リスト
- From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents [38.52713500119118]
Memoraは、数週間から数ヶ月のユーザ会話にまたがる長期メモリベンチマークです。
ベンチマークでは、記憶、推論、レコメンデーションの3つのメモリグラウンドタスクを評価している。
FAMA(Forgetting-Aware Memory Accuracy)は、古いメモリや無効メモリへの依存を罰するメトリクスである。
論文 参考訳(メタデータ) (2026-04-21T21:31:01Z) - AdaMem: Adaptive User-Centric Memory for Long-Horizon Dialogue Agents [49.63422082885992]
長軸対話エージェントのための適応型ユーザ中心メモリフレームワークであるAdaMemを提案する。
AdaMemは対話履歴をワーキング、エピソディック、ペルソナ、グラフメモリに整理する。
LoCoMo と PERSONAMEM ベンチマーク上での AdaMem の評価を行った。
論文 参考訳(メタデータ) (2026-03-17T13:22:54Z) - RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies [54.23445842621374]
記憶は、長い水平と歴史に依存したロボット操作にとって重要である。
近年,視覚言語アクション(VLA)モデルにメモリ機構が組み込まれ始めている。
本稿では,VLAモデルの評価と進展のための大規模標準ベンチマークであるRoboMMEを紹介する。
論文 参考訳(メタデータ) (2026-03-04T21:59:32Z) - REMem: Reasoning with Episodic Memory in Language Agent [32.63834745610879]
エピソードメモリを用いた構築と推論のためのフレームワークであるREMemについて述べる。
我々はREMemがMem0やHippoRAG 2のような時空間記憶システムよりも大幅に優れていることを示す。
REMemはまた、答えられない質問に対してより堅牢な拒絶行動を示す。
論文 参考訳(メタデータ) (2026-02-13T23:54:55Z) - Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents [76.76004970226485]
長期記憶はマルチモーダル大言語モデル(MLLM)エージェントにとって重要な機能である。
Mem-GalleryはMLLMエージェントのマルチモーダル長期会話メモリ評価のための新しいベンチマークである。
論文 参考訳(メタデータ) (2026-01-07T02:03:13Z) - WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning [66.24870234484668]
我々は,複数の相補的記憶から構築・取得する,新しいマルチモーダルメモリエージェント WorldMM を紹介する。
WorldMMは5つの長いビデオ質問回答ベンチマークで既存のベースラインを大幅に上回っている。
論文 参考訳(メタデータ) (2025-12-02T05:14:52Z) - Evaluating Long-Term Memory for Long-Context Question Answering [100.1267054069757]
質問応答タスクにアノテートした合成長文対話のベンチマークであるLoCoMoを用いて,メモリ拡張手法の体系的評価を行う。
以上の結果から,メモリ拡張アプローチによりトークン使用率が90%以上削減され,競争精度が向上した。
論文 参考訳(メタデータ) (2025-10-27T18:03:50Z) - Evaluating Long-Term Memory in 3D Mazes [10.224858246626171]
Memory Mazeはエージェントの長期記憶を評価するために設計されたランダム化迷路の3Dドメインである。
既存のベンチマークとは異なり、Memory Mazeはエージェントの能力から切り離された長期的なメモリを測定する。
現在のアルゴリズムは、時間の経過とともに縮小したバックプロパゲーションによるトレーニングの恩恵を受け、小さな迷路で成功するが、大きな迷路での人間のパフォーマンスに欠ける。
論文 参考訳(メタデータ) (2022-10-24T16:32:28Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。