論文の概要: Mnemon: Raw Records, Fast Judgments, Slow Thoughts
- arxiv url: http://arxiv.org/abs/2609.36059v1
- Date: Mon, 28 Sep 2026 18:17:13 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-30 21:28:46.931689
- Title: Mnemon: Raw Records, Fast Judgments, Slow Thoughts
- Title(参考訳): Mnemon:Raw Records、Fast Judgments、Slow Thoughts
- Abstract要約: メモリの作業は、考えるように、2つのシステムに分割する、と我々は主張する。
たいていは高速なシステム1の作業で、多くの小さな、独立したイエス/ノーのレコードに関する判断である。
この分割に基づいて構築されたメモリエージェントMnemonを提示する。
- 参考スコア(独自算出の注目度): 0.0
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Long-term memory lets an LLM assistant use a history it can no longer reread, and most memory systems build it by rewriting conversations into facts, graphs or typed memories at write time. We argue that the work of memory divides, as thinking does, into two systems. Most of it is fast System 1 work: many small, independent yes/no judgments about records, such as whether a record is needed or no longer current, which a decision model makes by the dozen in a third of a second. Only a little is slow System 2 work: writing a few search queries, naming what the reply needs and composing the answer, which an LLM does well but slowly. We present Mnemon, a memory agent built on this division. It keeps conversations as raw, dated records; an LLM (System 2) plans searches over them, a decision model, Jev (System 1), judges what the searches return, and rules with explicit budgets turn the judgments into a small View for an unchanged answering model. A background pass consolidates each record once into topic timelines, value histories and standing instructions linked to the records, so that questions about a whole conversation reach evidence their own searches miss. Because nothing is decided about a record when it is written, the same agent can read any store that returns dated records. With gpt-4.1-mini answering, as in a public re-evaluation of 14 systems, Mnemon scores 91.7% on LoCoMo, the highest among them, and 83.8% on LongMemEval-S, from under 4k tokens of context per question, with the lowest effective cost index on LoCoMo. With a reasoning model answering, it reaches 92.2% on LoCoMo and 94.4% on LongMemEval-S, the latter on par with the best published results. From 100K to 10M tokens of history on BEAM, its cost per question grows by a factor of 1.11. On the same records, Jev separates gold evidence better than two LLMs and is 3-11 times faster.
- Abstract(参考訳): 長期記憶により、LLMアシスタントは読み取り不能な履歴を使用でき、ほとんどのメモリシステムは、書き込み時に会話を事実、グラフ、型付きメモリに書き換えることでそれを構築できる。
メモリの作業は、考えるように、2つのシステムに分割する、と我々は主張する。
多くの場合、高速なシステム1の作業である: レコードが必要かどうか、もはや必要でないかどうかなど、レコードに関する小さな、独立したイエス/ノーな判断が、数十を3分の1秒に決定する決定モデルによって行われる。
システム2の動作が遅いのは、いくつかの検索クエリの記述、応答に必要なものの名前の指定、LLMがうまく機能するが、ゆっくりと動作する回答の作成です。
この分割に基づいて構築されたメモリエージェントMnemonを提示する。
LLM(System 2)はそれらを検索し、決定モデルであるJev(System 1)は、検索が返却するものを判断し、明確な予算を持つルールは、決定を変化しない応答モデルのための小さなビューに変換する。
バックグラウンドパスは、各レコードをトピックのタイムラインに集約し、履歴を値し、レコードにリンクしたスタンディングインストラクションを付加することで、会話全体に関する質問が、自身の検索ミスの証拠に到達する。
書き込み時にレコードについては何も決定されないため、同じエージェントが日付のレコードを返す任意のストアを読むことができる。
gpt-4.1-miniの回答では、14のシステムの再評価と同様に、Mnemonは、LoCoMoで91.7%、LongMemEval-Sで83.8%、質問毎に4k以下のコンテキストのトークンから、LoCoMoで最も効果的なコスト指数で評価されている。
推論モデルでは、LoCoMoで92.2%、LongMemEval-Sで94.4%に達する。
BEAM 上の 1 万から 10 万の履歴トークンまで、そのコストは 1.11 倍になる。
同じ記録では、Jevは2つのLDMよりも金のエビデンスを分離し、3.11倍高速である。
関連論文リスト
- CueMem: Cue-Guided Context Reconstruction for Long-Term Conversational Memory [38.76918926805768]
CueMemは、抽出されたメモリレコードを検索キューとして扱うキュー誘導フレームワークである。
ソースターンからクエリ関連対話コンテキストを再構築する。
LoCoMoとLongMemEvalの実験によると、CueMemは長期的なメモリベースラインを一貫して上回っている。
論文 参考訳(メタデータ) (2026-09-11T02:27:22Z) - Agent Zero Memory: Provenance-Aware Long-Term Memory for LLM Agents [3.6666917625778193]
Agent Zero Memory(エージェントゼロメモリ)は、長期記憶システムである。
ユーザの会話、ファイル、接続されたソースを3つの並列メモリシステムに消去する。
論文 参考訳(メタデータ) (2026-08-30T06:55:59Z) - RippleMem: From Isolated Retrieval to Associative Recollection for Long-Term Agent Memory [13.519130680921565]
本稿では,長期記憶システムであるRippleMemを紹介する。
LoCoMoとLongMemEval-Sは、RippleMemが評価された設定全体で最高の全体的なパフォーマンスを達成することを示している。
論文 参考訳(メタデータ) (2026-08-13T15:05:01Z) - Zero-Mem: Zero-Token Memory Operations for LLM Agents [17.228305588778]
Zero-Memは、元の相互作用トレースをレコードのソースとして保存する。
Zero-Memは、長期メモリと長期コンテキストの問合せベンチマークで競合するパフォーマンスを達成する。
論文 参考訳(メタデータ) (2026-07-31T13:01:06Z) - Less Context, More Accuracy: A Bi-Temporal Memory Engine for LLM Agents Where a Lean Retrieved Context Beats the Full History [2.5401434059780468]
Engramは、バイテンポラルデータモデル上のデュアルプロセスメモリエンジンである。
高速書き込みパスは、クリティカルパスにLSMなしでエピソードを付加する。
ハイブリッドリードパスは、密度、語彙、グラフ、および電流/セイレンス信号を融合する。
論文 参考訳(メタデータ) (2026-06-05T11:43:56Z) - MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models [50.25006399944962]
メモリは、長いマルチモーダル相互作用を扱うために、大きな視覚言語モデルにとって不可欠である。
MEMLENSはマルチモーダルマルチセッション会話におけるメモリのベンチマークである。
我々は27個のLVLMと7個のメモリ増強剤を評価した。
論文 参考訳(メタデータ) (2026-05-14T14:41:17Z) - From Single to Multi-Granularity: Toward Long-Term Memory Association and Selection of Conversational Agents [79.87304940020256]
大言語モデル(LLM)は会話エージェントで広く採用されている。
MemGASは、多粒度アソシエーション、適応選択、検索を構築することにより、メモリ統合を強化するフレームワークである。
4つの長期メモリベンチマークの実験により、MemGASは質問応答と検索タスクの両方において最先端の手法より優れていることが示された。
論文 参考訳(メタデータ) (2025-05-26T06:13:07Z) - Recursively Summarizing Enables Long-Term Dialogue Memory in Large Language Models [30.48902594738911]
長い会話をすると、大きな言語モデル(LLM)は過去の情報を思い出さず、一貫性のない応答を生成する傾向がある。
本稿では,長期記憶能力を高めるために,大規模言語モデル(LLM)を用いて要約/メモリを生成することを提案する。
論文 参考訳(メタデータ) (2023-08-29T04:59:53Z) - SCM: Enhancing Large Language Model with Self-Controlled Memory Framework [54.33686574304374]
大きな言語モデル(LLM)は、長い入力を処理できないため、重要な歴史的情報が失われる。
本稿では,LLMが長期記憶を維持し,関連する情報をリコールする能力を高めるための自己制御メモリ(SCM)フレームワークを提案する。
論文 参考訳(メタデータ) (2023-04-26T07:25:31Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。