論文の概要: Global Optimality for Constrained Exploration via Penalty Regularization
- arxiv url: http://arxiv.org/abs/2604.28144v1
- Date: Thu, 30 Apr 2026 17:31:46 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-05-01 16:31:54.225842
- Title: Global Optimality for Constrained Exploration via Penalty Regularization
- Title(参考訳): ペナルティ規則化による制約付き探査のグローバル最適性
- Authors: Florian Wolf, Ilyas Fatkhullin, Niao He,
- Abstract要約: 現実世界の探索は、しばしば客観的な制約によって制約された性質によって悪用される。
本稿では,一般凸擬ペナルティ正規化を実施するポリシグラディエント手法を提案する。
- 参考スコア(独自算出の注目度): 28.060709233558654
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Efficient exploration is a central problem in reinforcement learning and is often formalized as maximizing the entropy of the state-action occupancy measure. While unconstrained maximum-entropy exploration is relatively well understood, real-world exploration is often constrained by safety, resource, or imitation requirements. This constrained setting is particularly challenging because entropy maximization lacks additive structure, rendering Bellman-equation-based methods inapplicable. Moreover, scalable approaches require policy parameterization, inducing non-convexity in both the objective and the constraints. To our knowledge, the only prior model-free policy-gradient approach for this setting under general policy parameterization is due to Ying et al. (2025). Unfortunately, their guarantees are limited to weak regret and ergodic averages, which do not imply that the final output is a single deployable policy that is near-optimal and nearly feasible. In this work we take a different approach to this problem, and propose Policy Gradient Penalty (PGP) method, a single-loop policy-space method that enforces general convex occupancy-measure constraints via quadratic-penalty regularization. PGP constructs pseudo-rewards that yield gradient estimates of the penalized objective, subsequently exploiting the classical Policy Gradient Theorem. We further establish the regularity of the penalized objective, providing the smoothness properties needed to justify the convergence of PGP. Leveraging hidden convexity and strong duality, we then establish global last-iterate convergence guarantees, attaining an $ε$-optimal constrained entropy value with $ε$ bounded constraint violation despite policy-induced non-convexity. We validate PGP through ablations on a grid-world benchmark and further demonstrate scalability on two challenging continuous-control tasks.
- Abstract(参考訳): 効率的な探索は強化学習の中心的な問題であり、しばしば国家行動占有率のエントロピーを最大化するものとして形式化される。
制限のない最大エントロピー探索は比較的よく理解されているが、現実世界の探索は安全、資源、模倣の要求によって制約されることが多い。
エントロピーの最大化は加法構造を欠いているため、ベルマン方程式に基づく手法を適用できないため、この制約された設定は特に困難である。
さらに、スケーラブルなアプローチはポリシーのパラメータ化を必要とし、目的と制約の両方において非凸性を引き起こす。
我々の知る限り、一般的な政策パラメータ化の下でのこの設定に対する事前のモデルフリー政策段階的アプローチは、Ying et al (2025) によるものである。
残念なことに、彼らの保証は弱い後悔とエルゴード的な平均に限られており、最終的な出力は、ほぼ最適でほぼ実現可能な単一のデプロイ可能なポリシーであることを意味するわけではない。
本研究では,この問題に対して異なるアプローチをとるとともに,一般凸占有制約を2次ペナルティ正規化を通じて適用する単一ループポリシー空間法であるポリシグラディエントペナルティ(PGP)法を提案する。
PGPは、ペナル化対象の勾配推定を導出する擬逆数を構築し、その後古典的なポリシー勾配定理を利用する。
さらに、PGPの収束を正当化するために必要な滑らか性特性を提供する。
隠れ凸性と強い双対性を利用することで、政策による非凸性にもかかわらず、$ε$最適制約エントロピー値が$ε$有界制約違反となる大域的最終点収束を保証する。
グリッドワールドベンチマークの短縮を通じてPGPを検証するとともに、2つの困難な連続制御タスクのスケーラビリティをさらに実証する。
関連論文リスト
- Last-Iterate Global Convergence of Policy Gradients for Constrained Reinforcement Learning [62.81324245896717]
我々はC-PGと呼ばれる探索非依存のアルゴリズムを導入し、このアルゴリズムは(弱)勾配支配仮定の下でのグローバルな最終点収束を保証する。
制約付き制御問題に対して,我々のアルゴリズムを数値的に検証し,それらを最先端のベースラインと比較する。
論文 参考訳(メタデータ) (2024-07-15T14:54:57Z) - Last-Iterate Convergent Policy Gradient Primal-Dual Methods for
Constrained MDPs [107.28031292946774]
無限水平割引マルコフ決定過程(拘束型MDP)の最適ポリシの計算問題について検討する。
我々は, 最適制約付きポリシーに反復的に対応し, 非漸近収束性を持つ2つの単一スケールポリシーに基づく原始双対アルゴリズムを開発した。
我々の知る限り、この研究は制約付きMDPにおける単一時間スケールアルゴリズムの非漸近的な最後の収束結果となる。
論文 参考訳(メタデータ) (2023-06-20T17:27:31Z) - Convergence and sample complexity of natural policy gradient primal-dual methods for constrained MDPs [23.596546979904613]
線形決定(MDP)の割引最適率の解法として, 自然政策勾配法を用いる。
また、2つのサンプルベースNPG-PDアルゴリズムに対して有限サンプル保証を提供する。
論文 参考訳(メタデータ) (2022-06-06T04:28:04Z) - Beyond Exact Gradients: Convergence of Stochastic Soft-Max Policy Gradient Methods with Entropy Regularization [20.651913793555163]
古典的エントロピー正規化政策勾配法をソフトマックス政策パラメトリゼーションで再検討する。
提案したアルゴリズムに対して,大域的最適収束結果と$widetildemathcalO(frac1epsilon2)$のサンプル複雑性を確立する。
論文 参考訳(メタデータ) (2021-10-19T17:21:09Z) - Risk-Sensitive Deep RL: Variance-Constrained Actor-Critic Provably Finds
Globally Optimal Policy [95.98698822755227]
本研究は,リスクに敏感な深層強化学習を,分散リスク基準による平均報酬条件下で研究する試みである。
本稿では,ポリシー,ラグランジュ乗算器,フェンシェル双対変数を反復的かつ効率的に更新するアクタ批判アルゴリズムを提案する。
論文 参考訳(メタデータ) (2020-12-28T05:02:26Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。