Fugu-MT 論文翻訳(概要): HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime

論文の概要: HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime

arxiv url: http://arxiv.org/abs/2605.30201v1
Date: Thu, 28 May 2026 16:38:21 GMT
ステータス: 翻訳完了
システム内更新日: 2026-05-30 02:45:56.54209
Title: HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime
Title（参考訳）: HPO:スパルス・リワード・レジーム下における安定・効率的なトレーニングのためのヒステリックポリシー最適化
Authors: Mohamed Sana, Nicola Piovesan, Antonio De Domenico, Fadhel Ayed, Haozhe Zhang,
Abstract要約: Hysteretic Policy Optimization (HPO)は、負のアドバンテージ更新の重みを減らす。 Adaptive HPOはバッチレベルのアドバンテージサイン統計に基づいてヒステリックウェイトを設定する。 TeleLogsでは、A-HPOが0.84で、SAPOを5%上回り、GSPOを11%上回り、GRPOを15%上回る。
参考スコア（独自算出の注目度）: 7.472260601349898
License: http://creativecommons.org/licenses/by/4.0/
Abstract: We investigate a narrow but common failure mode of GRPO-style reinforcement learning in the context of sparse verifiable rewards: early updates contain more responses with negative advantages than those with positive advantages, while response-level length normalization ties the magnitude of the update to the length of the output. We propose Hysteretic Policy Optimization (HPO), a minimal modification of GRPO that reduces the weight of negative-advantage updates and replaces per-response length normalization with mean-length normalization. We further introduce Adaptive HPO (A-HPO), which sets the hysteretic weight based on batch-level advantage-sign statistics, thereby removing the need for tuning a fixed hysteretic weight. In our TeleLogs and Countdown experiments, A-HPO improves the reward per update compared to GRPO, with the largest gains in early sparse reward regimes. On TeleLogs, A-HPO achieves a final reward of 0.84, outperforming SAPO by 5%, GSPO by 11%, and GRPO by 15%, while maintaining a comparable response-length. On Countdown, A-HPO achieves the largest gains in initial and most difficult configurations across 1.5B-7B models. Ablation studies on the hysteretic weight show that the gains of A-HPO come from better balancing the contributions of positive and negative advantages compared to positive-only or fully symmetric updates.
Abstract（参考訳）: 初期更新は、正の利点を持つものよりも負の利点を持つものが多く、応答レベル長正規化は、出力の長さに対する更新の規模に結びついている。負のアドバンテージ更新の重みを減らし,応答長ごとの正規化を平均長正規化に置き換える,GRPOの最小限の修正であるHysteretic Policy Optimization (HPO)を提案する。さらに、バッチレベルの利点符号統計に基づいてヒステリックウェイトを設定するアダプティブHPO(A-HPO)を導入し、固定ヒステリックウェイトを調整する必要をなくした。 TeleLogsとCountdownの実験では、A-HPOはGRPOと比較してアップデート当たりの報酬を改善しています。 TeleLogsでは、A-HPOが0.84で、SAPOを5%上回り、GSPOを11%上回り、GRPOを15%上回り、同等の応答長を維持している。 Countdownでは、A-HPOは1.5B-7Bモデルで初期および最も難しい構成で最大のゲインを達成している。ヒステリックウェイトに関するアブレーション研究は、A-HPOの利得は、正にのみあるいは完全に対称な更新に比べて、正と負の利点の寄与のバランスが良くなることを示している。

論文の概要: HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime

関連論文リスト