論文の概要: ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams
- arxiv url: http://arxiv.org/abs/2607.09759v2
- Date: Tue, 14 Jul 2026 07:18:26 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-15 12:51:44.50856
- Title: ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams
- Title(参考訳): ReflectWorld-MM: オープンエンディングビデオストリームのためのエンティティ指向マルチモーダルメモリシステム
- Authors: Xiaokang Ma, Yifan Sun, Zhihong Jin, Jie Gu, Yudong Luo, Shenyi Shao, Chu Tang, Jingmin Chen, Li Pu,
- Abstract要約: オープンなビデオストリームのためのエンティティ指向マルチモーダルメモリシステムであるReflectWorld-MMを提案する。
これは3つの部分から構成されており、第1の部分は知覚フロントエンドであり、オーディオ視覚ストリームを、境界付き短期記憶下での実体分解された観察に変換する。
2つ目は階層的な長期記憶であり、人間の記憶理論に基礎を置いており、多スケールのエピソード記憶、進化するエンティティ中心のセマンティックメモリ、手続き記憶を結合している。
- 参考スコア(独自算出の注目度): 8.099856183300002
- License: http://creativecommons.org/licenses/by-nc-nd/4.0/
- Abstract: Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-standing goal, and recently multimodal agents equipped with long-term memory over video streams have attracted increasing interest. Unfortunately, existing systems either keep their memory inside the model context or in a flat feature store, and organize it around frames rather than around the persistent entities a stream is really about, which confines them to bounded videos and weakens their ability to track who and what reappears over time. In this paper, we propose ReflectWorld-MM, an entity-oriented multimodal memory system for open-ended video streams. It consists of three parts. The first is a perception front-end that turns an audiovisual stream into entity-resolved observations under a bounded short-term memory. The second is a hierarchical long-term memory, grounded in human memory theory, that couples a multi-scale episodic memory, an evolving entity-centric semantic memory, and a procedural memory. The third is a complete realization, built for real-world operation, that ingests arbitrary streams and plugs into off-the-shelf assistants. Across six long-video and lifelong-memory benchmarks, ReflectWorld-MM achieves the best accuracy on all six, outperforming strong memory agents and a frontier model.
- Abstract(参考訳): 世界を継続的に監視し、見ているものを思い出し、蓄積した経験を理由づけるアシスタントを構築することは長年の目標であり、近年はビデオストリームに長期記憶を備えたマルチモーダルエージェントが関心を集めている。
残念なことに、既存のシステムは、メモリをモデルコンテキスト内に保持するか、フラットなフィーチャーストア内に保持し、ストリームが本当に意味する永続的なエンティティではなく、フレームの周りに配置する。
本稿では,オープンエンドビデオストリームのためのエンティティ指向マルチモーダルメモリシステムであるReflectWorld-MMを提案する。
3つの部分から構成される。
1つ目は、視覚的ストリームを、境界付き短期記憶下での実体分解された観察に変換する知覚フロントエンドである。
2つ目は階層的な長期記憶であり、人間の記憶理論に基礎を置いており、多スケールのエピソード記憶、進化するエンティティ中心のセマンティックメモリ、手続き記憶を結合している。
3つ目は、現実の操作のために作られた完全な実現で、任意のストリームを取り込み、既製のアシスタントにプラグインする。
リフレクションワールド-MMは6つの長ビデオおよび寿命メモリのベンチマークで6つすべてで最高の精度を達成し、強力なメモリエージェントとフロンティアモデルを上回っている。
関連論文リスト
- Homer: Understanding Long-form Videos with Hierarchical Memory and Agentic Reasoning [77.7531245219403]
textscHomerは階層型オンラインメモリ探索および推論フレームワークである。
textscHomerの記憶は、生の知覚から繰り返される実体、明示的な時間的・因果関係と結びついた出来事まで、長いビデオのマルチスケール構造を反映している。
textscHomerは、M3-Bench-robot、M3-Bench-web、Video-MME-Longで$+5.5$、$+10.8$、$+4.4$ポイントで前のベストエージェントメソッドより優れている。
論文 参考訳(メタデータ) (2026-07-01T05:53:09Z) - WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction [72.1620416874118]
マルチモーダルな言語モデルは、長距離エージェントとしてますます多くデプロイされている。
既存のベンチマークは、静的対話上のリコールを測定し、メモリを1つのタスクの精度に分解し、キャプションに対する視覚的な観察を減らす。
マルチモーダルエージェントメモリを,観測可能な4段階ライフサイクルを持つアクションワールドインタラクションループとして定式化する。
論文 参考訳(メタデータ) (2026-05-28T04:27:20Z) - EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding [89.26501160264199]
EgoMemReasonは、メモリ駆動推論を通じて、1週間のエゴセントリックなビデオ理解を体系的に評価する。
EgoMemReasonには3つのメモリタイプと6つのコア課題に関する500の質問が含まれている。
EgoMemReasonをMLLMとエージェントフレームワークにまたがる17の手法で評価する。
論文 参考訳(メタデータ) (2026-05-11T01:59:59Z) - From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents [78.30630000529133]
本稿ではファジィトレース理論に基づくピラミッド型マルチモーダルメモリアーキテクチャMM-Memを提案する。
MM-Memメモリは階層的に感覚バッファ、エピソードストリーム、シンボリックに構造する。
実験により、MM-Memがオフラインタスクとストリーミングタスクの両方で有効であることが確認された。
論文 参考訳(メタデータ) (2026-03-02T05:12:45Z) - WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning [66.24870234484668]
我々は,複数の相補的記憶から構築・取得する,新しいマルチモーダルメモリエージェント WorldMM を紹介する。
WorldMMは5つの長いビデオ質問回答ベンチマークで既存のベースラインを大幅に上回っている。
論文 参考訳(メタデータ) (2025-12-02T05:14:52Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。