論文の概要: Coding Agent Memory Post-training: Unlocking the Memory Potential of Pre-trained File Operations for Long-Horizon Tasks via Reinforcement Learning
- arxiv url: http://arxiv.org/abs/2609.34422v2
- Date: Wed, 30 Sep 2026 13:09:33 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-01 18:57:25.999684
- Title: Coding Agent Memory Post-training: Unlocking the Memory Potential of Pre-trained File Operations for Long-Horizon Tasks via Reinforcement Learning
- Title(参考訳): 符号化エージェントのメモリ後学習:強化学習による長軸タスクのための事前学習ファイル操作のメモリポテンシャルの解錠
- Abstract要約: 本稿では,CAMG,Shop,Coding,DeepResearch,AutoResearchにまたがる長距離エージェント-RL環境について紹介する。
CAMGは実行可能なシェルアクセスとエピソードパーシスタントなワークスペースを提供し、エージェントがファイルのメモリとして作成、修正、検索、再利用を可能にする。
また,CAMG-RLを導入し,完全に非同期なPPOを持つ4つの環境すべてに対して,単一のポリシを共同でトレーニングする。
- 参考スコア(独自算出の注目度): 23.27669373387788
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Language-model agents increasingly tackle long-horizon tasks whose interaction histories exceed the model's active context. Recent work has begun to use reinforcement learning to make memory control part of the policy, often relying on predefined memory tools within domain-specific training environments of relatively short horizons. This setup ties learned memory behavior to environment-specific interfaces that lie outside the base model's pre-training and must be learned from scratch, so even after post-training, agents struggle to use memory in long-horizon tasks. To address these limitations, we introduce Coding Agent Memory Gym (CAMG), a suite of long-horizon agentic-RL environments spanning Shop, Coding, DeepResearch, and AutoResearch. Alongside each environment's native task interface, CAMG provides executable shell access and an episode-persistent workspace, enabling agents to create, revise, search, and reuse files as memory throughout an episode. We also introduce CAMG-RL, which trains a single policy jointly across all four environments with fully asynchronous PPO, learning this file-based memory behavior directly from downstream task reward, and we train CAMG-RL-4B and CAMG-RL-9B from Qwen3.5 models of matching size. On SWE-bench Verified and MLE-bench Lite, CAMG-RL-4B and CAMG-RL-9B are competitive with Qwen3.5-35B-A3B and Qwen3.5-122B-A10B, respectively.
- Abstract(参考訳): 言語モデルエージェントは、相互作用履歴がモデルのアクティブなコンテキストを超えた長い水平タスクにますます取り組みます。
最近の研究では、比較的短い地平線のドメイン固有のトレーニング環境において、メモリ管理を予め定義されたメモリツールに依存して、強化学習をポリシーの一部にし始めている。
このセットアップは、ベースモデルの事前トレーニングの外にあり、スクラッチから学ぶ必要がある環境固有のインターフェースにメモリの振る舞いを学習する。
これらの制限に対処するため、Coding Agent Memory Gym (CAMG)を導入し、Shop、Coding、DeepResearch、AutoResearchにまたがる長距離エージェントRL環境のスイートを紹介した。
それぞれの環境のネイティブタスクインターフェースに加えて、CAMGは実行可能シェルアクセスとエピソード永続化ワークスペースを提供し、エージェントはエピソードを通してファイルを作成し、修正し、検索し、再利用することができる。
また,CAMG-RLは,完全非同期PPOで4つの環境すべてに一貫した単一ポリシをトレーニングし,このファイルベースのメモリ動作を下流タスク報酬から直接学習し,マッチングサイズをQwen3.5モデルからCAMG-RL-4BとCAMG-RL-9Bを訓練する。
SWE-bench VerifiedとMLE-bench Liteでは、CAMG-RL-4BとCAMG-RL-9BはそれぞれQwen3.5-35B-A3BとQwen3.5-122B-A10Bと競合する。
関連論文リスト
- Dual-Grained Agent Memory and Shapley Context Attribution for Multimodal Agentic Learner [55.80248281761638]
マルチモーダルな大言語モデル (MLLM) は、科学的および数学的推論において、印象的な認識を提供する。
本稿では,DG-Memを提案する。DG-Memは,トレーニングタイムのロールアウトから一度構築した外部記憶メモリと,テスト時に読み取り専用に相談することで,凍結したMLLMを拡張可能なエージェントメモリフレームワークである。
論文 参考訳(メタデータ) (2026-08-24T13:56:01Z) - MemGym: a Long-Horizon Memory Environment for LLM Agents [69.79226770543049]
本稿では,エージェントメモリのベンチマークであるMemGymを紹介する。
MemGymは、メモリパフォーマンスを推論、検索、ツール使用能力から切り離すメモリアイソレーションスコアを報告している。
MEMGYM-CODEQAとMEMGYM-DRの合成パイプラインは、長さ制御可能であり、各ステージでアブレーションを検証可能であり、下流のシナリオと密に整合している。
論文 参考訳(メタデータ) (2026-05-20T07:25:33Z) - Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning [89.55738101744657]
大規模言語モデル(LLM)は、幅広いNLPタスクで印象的な機能を示しているが、基本的にはステートレスである。
本稿では,LLMに外部メモリを積極的に管理・活用する機能を備えた強化学習フレームワークであるMemory-R1を提案する。
論文 参考訳(メタデータ) (2025-08-27T12:26:55Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。