論文の概要: Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning
- arxiv url: http://arxiv.org/abs/2608.05111v1
- Date: Wed, 05 Aug 2026 17:44:52 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-06 14:48:44.065264
- Title: Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning
- Title(参考訳): 強化学習におけるエピソード探索とニューラルメモリの相互作用の逆構造形成
- Authors: Jai Malegaonkar, Rohan Patil, Henrik I. Christensen,
- Abstract要約: 強化学習では、エージェントは報酬のある状態に遭遇し、その経験を記憶に残して、彼らのポリシーを最適化しなければならない。
本稿では,多種多様な記憶アーキテクチャを用いた横断的探索ボーナスについて検討する。
我々の結果は、探索と記憶は補足であり、代替物ではないことを示している。
- 参考スコア(独自算出の注目度): 2.1104538070656087
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: In partially observable reinforcement learning, agents face a dual bottleneck: they must explore to encounter rewarding states and retain that experience in memory to optimize their policies. Exploration bonuses and memory architectures are traditionally evaluated in isolation, leaving their interaction unmeasured, and standard notions of sparse reward conflate temporal signal density with what the reward actually supervises. We present a controlled study crossing episodic exploration bonuses with diverse neural memory architectures across three environments that vary how the content of memory is acquired. An identical bonus signal yields three distinct interaction patterns: it amplifies architectural capacity differences where memory content must be actively discovered and retained unsupervised; equalizes architectures to a shared ceiling where the content, once sought out, is a single reward-supervised cue; and is null where the observation stream is purely scheduled. Controlled reward manipulations verify that these patterns track reward structure rather than density: a dense reward neutralizes a bonus only if it directly supervises the required latent memory, and a small avoidable penalty on exploratory actions (leaving the optimum unchanged) induces policy convergence to suboptimal stationary states, which either bonus resolves. We then formalize reward sparsity with observation-anchored reward machines, separating structural sparsity (an automaton reproduces the return without the task-required history) from potential sparsity (the one-step reward misprices local exploratory actions); the resulting vocabulary organizes the three regimes by the retention burden each task exposes. Together, these results show exploration and memory are complements, not substitutes: a bonus induces exposure, and only memory converts exposure into return.
- Abstract(参考訳): エージェントは報酬のある状態に遭遇し、その経験を記憶に残して、ポリシーを最適化しなければならない。
探索ボーナスとメモリアーキテクチャは、伝統的に独立して評価され、その相互作用は測定されず、スパース報酬の標準的な概念は、時間信号密度と、その報酬が実際に監督するものとを区別する。
本稿では,記憶内容の獲得方法の異なる3つの環境にまたがる多種多様な記憶アーキテクチャを用いたエピソード探索ボーナスについて検討する。
同一のボーナス信号は、3つの異なる相互作用パターンを生成する: メモリ内容がアクティブに発見され、管理されていないまま保持されるアーキテクチャ容量の差異を増幅する; アーキテクチャを、一度検索されたコンテンツが単一の報酬管理キューである共有天井に等化する; 観測ストリームが純粋にスケジュールされているヌルである。
制御された報酬操作は、これらのパターンが密度よりも報酬構造を追跡することを検証する: 密度の高い報酬は、要求された潜伏メモリを直接監督する場合にのみボーナスを中和し、探索的行動(最適に変化しないまま)に対する小さな回避可能なペナルティは、ボーナスが解決する準最適定常状態へのポリシー収束を誘導する。
次に,各課題が露呈する保持負担によって3つの体制を整理し,構造的疎大性(タスク要求履歴無しでリターンを再現するオートマトン)を潜在的疎大性(一段階の報酬が局所探索行動に悪影響を及ぼす)から分離する。
ボーナスは露光を誘導し、メモリだけが露光を返却する。
関連論文リスト
- Beyond Binary Rewards: A Comparative Study of Reward Design for Reinforcement Unlearning [18.91502176461499]
機械学習は、訓練された言語モデルからの特定の知識を、完全なリトレーニングなしで取り除こうとする。
最近の研究は、未学習をReinforcement Learning with Verifiable Rewards (RLVR)問題として再考している。
報奨設計が強化アンラーニングフレームワークにおける非学習効率に与える影響について検討する。
論文 参考訳(メタデータ) (2026-07-30T10:15:59Z) - SaliMory: Orchestrating Cognitive Memory for Conversational Agents [22.813290545851743]
SALIMORYは、認知的に構造化されたメモリスパンニングされたユーザファクト、好み、ワーキングメモリを管理するために単一の言語モデルをトレーニングするフレームワークである。
メモリ分散障害を3分の1削減し、エンドツーエンドの精度で最先端を10%以上上回り、良質なパーソナライゼーション率を2倍以上に向上させる。
論文 参考訳(メタデータ) (2026-06-02T18:31:50Z) - Evoking User Memory: Personalizing LLM via Recollection-Familiarity Adaptive Retrieval [59.295767860331004]
RF-Memは、親しみやすい不確実性誘導デュアルパスメモリレトリバーである。
それは、人間のようなデュアルプロセス認識をレトリバーに埋め込む。
一定の予算とレイテンシの制約の下で、ワンショット検索とフルコンテキスト推論を一貫して上回る。
論文 参考訳(メタデータ) (2026-03-10T06:31:44Z) - Fine-Mem: Fine-Grained Feedback Alignment for Long-Horizon Memory Management [63.48041801851891]
Fine-Memは、きめ細かいフィードバックアライメントのために設計された統一されたフレームワークである。
MemalphaとMemoryAgentBenchの実験は、Fin-Memが強いベースラインを一貫して上回ることを示した。
論文 参考訳(メタデータ) (2026-01-13T11:06:17Z) - Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution [52.76038908826961]
我々は静的ストレージと動的推論のギャップを埋めるため、$textbfReMe$ ($textitRemember Me, Refine Me$)を提案する。
ReMeは3つのメカニズムを通じてメモリライフサイクルを革新する: $textitmulti-faceted distillation$, きめ細かい経験を抽出する。
BFCL-V3とAppWorldの実験では、ReMeが新しい最先端のエージェントメモリシステムを確立している。
論文 参考訳(メタデータ) (2025-12-11T14:40:01Z) - VOGUE: Guiding Exploration with Visual Uncertainty Improves Multimodal Reasoning [62.09195763860549]
検証可能な報酬(RLVR)による強化学習は、大きな言語モデル(LLM)の推論を改善するが、探索に苦労する。
出力(テキスト)から入力(視覚)空間へ探索をシフトする新しい手法である$textbfVOGUE(Visual Uncertainty Guided Exploration)を紹介した。
本研究は,視覚入力の本質的不確実性における基盤探索が,マルチモーダル推論を改善するための効果的な戦略であることを示す。
論文 参考訳(メタデータ) (2025-10-01T20:32:08Z) - Towards Better De-raining Generalization via Rainy Characteristics Memorization and Replay [74.54047495424618]
現在の画像のデライニング方法は、主に限られたデータセットから学習する。
ネットワークが段階的にデライニングの知識基盤を拡大することを可能にする新しいフレームワークを導入する。
論文 参考訳(メタデータ) (2025-06-03T05:50:00Z) - Towards Control-Centric Representations in Reinforcement Learning from
Images [43.21376020617722]
ReBisは、報酬なしの制御情報と報酬特化知識を統合することで、制御中心の情報を取得することを目指している。
AtariゲームとDeepMind Control Suitを含む2つの大規模なベンチマークに関する実証研究は、ReBisが既存の方法よりも優れたパフォーマンスを示している。
論文 参考訳(メタデータ) (2023-10-25T14:09:53Z) - Successor-Predecessor Intrinsic Exploration [18.440869985362998]
本研究は,内因性報酬を用いた探索に焦点を当て,エージェントが自己生成型内因性報酬を用いて外因性報酬を過渡的に増強する。
本研究では,先進情報と振り返り情報を組み合わせた新たな固有報酬に基づく探索アルゴリズムSPIEを提案する。
本研究は,SPIEが競合する手法よりも少ない報酬とボトルネック状態の環境において,より効率的かつ倫理的に妥当な探索行動をもたらすことを示す。
論文 参考訳(メタデータ) (2023-05-24T16:02:51Z) - Automatic Recall Machines: Internal Replay, Continual Learning and the
Brain [104.38824285741248]
ニューラルネットワークのリプレイには、記憶されたサンプルを使ってシーケンシャルなデータのトレーニングが含まれる。
本研究では,これらの補助サンプルをフライ時に生成する手法を提案する。
代わりに、評価されたモデル自体内の学習したサンプルの暗黙の記憶が利用されます。
論文 参考訳(メタデータ) (2020-06-22T15:07:06Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。