論文の概要: SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue
- arxiv url: http://arxiv.org/abs/2609.26780v2
- Date: Wed, 23 Sep 2026 12:35:20 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-25 00:05:17.68062
- Title: SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue
- Title(参考訳): SpeakerMem-R1: 多人数対話のための話者中心デュアルトラックメモリ
- Abstract要約: SpeakerMem-R1は、メッセージと状態を個人レベルのビューとグループレベルのビューに整理したデュアルトラックメモリである。
SpeakerMem-R1は、それぞれ47.9%、69.2%、61.9%のバイナリアキュラシーを達成している。
また、1,986のLoCoMo質問に対して70.85%を達成し、2対1の会話境界テストとして使用しています。
- 参考スコア(独自算出の注目度): 30.06285587936361
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Long-term conversational memory in multi-party settings requires more than retrieving relevant content from long-term conversations: it must distinguish who said what, whom each statement concerns, how individuals perceive one another, what information is shared by the group, and how states change over time. Recent studies on multi-party dialogue benchmarks show that existing general-purpose LLM memory systems tend to lose person and group relations or struggle to integrate clues distributed across members, groups, and time. Together, these issues reveal two core bottlenecks: message attribution and relational understanding in multi-party dialogue, and state reconstruction from interleaved histories. To address both, we propose $\textbf{SpeakerMem-R1}$: its dual-track memory stores speaker-labeled verbatim messages and derived states organized into person-level and group-level views, then combines evidence from both tracks by entity, event, and time at query time. To reduce attribution and update errors during structured memory construction while enabling local deployment, we train Writer-R1 with SpeakerLevenshtein and speaker-conditioned GRPO. On GroupMemBench, SocialMemBench, and EverMemBench, SpeakerMem-R1 achieves binary accuracies of 47.9%, 69.2%, and 61.9%, respectively. On the publicly reported EverMemBench leaderboard from EverMind-AI, we achieves 62.33%, the best reported result among the latest state-of-the-art frameworks. It also achieves 70.85% on all 1,986 LoCoMo questions, which we use as a two-person long-term conversation boundary test. In a controlled evaluation of 305 questions, RL raises the SFT Writer's mean accuracy from 57.38% to 68.20%. We report both binary accuracy and token-F1, and ablations show that the verbatim and structured tracks, as well as person-level and group-level views, are complementary under the standardized evaluation interface.
- Abstract(参考訳): 長期的な会話記憶は、長期的な会話から関連したコンテンツを抽出する以上のものを必要としている:誰、誰がどの発言に関心を持つか、個人がお互いをどう知覚するか、どの情報がグループによって共有されているか、そして国家が時間とともにどのように変化するか、を区別する必要がある。
近年のマルチパーティ・ダイアログ・ベンチマークでは、既存の汎用LDMメモリシステムは人やグループの関係をなくしたり、メンバー、グループ、時間に分散した手がかりを統合するのに苦労する傾向にある。
これらの課題は、多人数対話におけるメッセージ属性と関係理解の2つの中心的ボトルネックと、インターリーブされた歴史からの状態復元である。
この2つのトラックメモリは、話者ラベル付き冗長メッセージと、個人レベルおよびグループレベルのビューに編成された派生状態を格納し、エンティティ、イベント、クエリ時に両方のトラックからのエビデンスを結合する。
ローカルな配置を可能にしながら、構造化メモリ構築時の属性の低減とエラーの更新を行うため、SpeakerLevenshteinと話者条件GRPOを用いてWriter-R1を訓練する。
GroupMemBench、SocialMemBench、EverMemBenchでは、SpeakerMem-R1はそれぞれ47.9%、69.2%、61.9%のバイナリアキュラシーを達成している。
EverMind-AIから公開されたEverMemBenchのリーダーボードで、最新の最先端フレームワークの中で最も報告された結果である62.33%を達成した。
また、1,986のLoCoMo質問に対して70.85%を達成し、2対1の会話境界テストとして使用しています。
305質問の制御された評価では、RL は SFT Writer の平均精度を 57.38% から 68.20% に引き上げる。
本稿では,2値精度とトークンF1の両方を報告し,評価インタフェースの標準化により,動詞と構造化されたトラックと,個人レベルのビューとグループレベルのビューが相補的であることを示す。
関連論文リスト
- EM^2Mem: Event-Centric Multimodal Memory for Large Language Models [52.3165591032625]
本稿では,イベント中心型マルチモーダルメモリフレームワークEM2Memを提案する。
3つの長ビデオQAベンチマークで、EM2Memは最強のメモリベースラインの平均精度を2.0、2.4、および3.7ポイント改善した。
論文 参考訳(メタデータ) (2026-09-01T01:38:41Z) - GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations [25.703133924514884]
大規模言語モデル(LLM)エージェントは、ますますパーソナルアシスタントや職場の協力者として機能している。
既存のメモリシステムとベンチマークは、Dyadicのシングルユーザ設定を中心に構築されている。
グループメモリの3つの特性を公開するベンチマークであるGroupMemBenchを紹介する。
論文 参考訳(メタデータ) (2026-05-14T07:38:29Z) - EverMemBench: Benchmarking Long-Term Interactive Memory in Large Language Models [16.865998112859604]
EverMemBenchは、100万以上のトークンにまたがる多人数のマルチグループ会話を特徴とするベンチマークである。
EverMemBenchは、1000以上のQAペアを通じて3次元にわたるメモリシステムを評価する。
論文 参考訳(メタデータ) (2026-02-01T16:13:08Z) - Memory-T1: Reinforcement Learning for Temporal Reasoning in Multi-session Agents [80.33280979339123]
強化学習(RL)を用いた時間認識メモリ選択ポリシーを学習するフレームワークであるMemory-T1を紹介する。
Time-Dialogベンチマークでは、Memory-T1が7Bモデルを67.0%に引き上げ、オープンソースモデルの新たな最先端パフォーマンスを確立した。
論文 参考訳(メタデータ) (2025-12-23T06:37:29Z) - A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results [62.01871490859886]
第9回CHiMEチャレンジにおいて,マルチモーダルコンテキスト認識(MCoRec)の課題を紹介した。
MCoRecは、録音が説明のない、カジュアルなグループチャットに集中する、自然なマルチパーティの会話をキャプチャする。
このタスクでは、各話者のスピーチを共同で翻訳し、音声・視覚録音から各話者の会話にまとめることにより、「誰がいつ、何、誰と話をするのか?」という質問に答えるシステムが必要である。
論文 参考訳(メタデータ) (2025-10-27T12:36:43Z) - On Memory Construction and Retrieval for Personalized Conversational Agents [69.46887405020186]
本稿では,セグメンテーションモデルを導入し,セグメントレベルでメモリバンクを構築するセグメンテーション手法であるSeComを提案する。
実験結果から,SeComは長期会話ベンチマークLOCOMOとLong-MT-Bench+のベースラインよりも優れた性能を示した。
論文 参考訳(メタデータ) (2025-02-08T14:28:36Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。