論文の概要: Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
- arxiv url: http://arxiv.org/abs/2605.09214v1
- Date: Sat, 09 May 2026 23:17:46 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-05-12 23:28:50.117555
- Title: Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
- Title(参考訳): 前向きKL規則化によるオフライン帯域の単一ポリシング性を考慮した高速速度
- Abstract要約: emphKullback-Leibler (KL) 正規化は強化学習アルゴリズムにおいてユビキタスである。
近年の研究では、KLの逆正則化の下での意思決定において、$1$型高速速度が示されている。
我々は、この問題を解決するための第一歩として、フォワードKL正規化オフラインCBの合理化分析を行う。
- 参考スコア(独自算出の注目度): 54.40598524756038
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: \emph{Kullback-Leibler} (KL) regularization is ubiquitous in reinforcement learning algorithms in the form of \emph{reverse} or \emph{forward} KL. Recent studies have demonstrated $ε^{-1}$-type fast rates for decision making under reverse KL regularization, in contrast to the standard $ε^{-2}$-type sample complexity. However, for forward-KL-regularized objectives, existing statistical analyses are either not applicable or result in $\tilde{O}(ε^{-2})$ slow rates. We take the first step towards addressing this problem via a streamlined analysis of forward-KL-regularized offline CBs. We give the first $\tilde{O}(ε^{-1})$ upper bounds in tabular and general function approximation settings, both under notions of \emph{single-policy concentrability}. In particular, our convex-analytical pipeline unifies these settings by exploiting the pessimism principle in a novel way and completely bypasses the proof routines in previous works based on the mean value theorem, which might be of independent interest. Moreover, we provide rate-optimal lower bounds, manifesting the tightness of our upper bounds in terms of statistical rates. Our lower bounds also demonstrate that the forward-KL-regularized sample complexity recovers the unregularized slow rate in the low-regularization regime, similarly to the reverse-KL regularization.
- Abstract(参考訳): \emph{Kullback-Leibler} (KL) 正規化は、強化学習アルゴリズムにおいて \emph{reverse} または \emph{forward} KL の形でユビキタスである。
最近の研究では、標準の$ε^{-2}$-typeサンプルの複雑さとは対照的に、逆KL正規化の下での意思決定のための$ε^{-1}$-type fast rateが示されている。
しかし、前方KL正規化の目的に対しては、既存の統計分析は適用できないか、あるいは$\tilde{O}(ε^{-2})$ slow rateとなる。
我々は,この問題を解決するための第一歩として,フォワードKL規則化オフラインCBの合理化分析を行う。
最初の$\tilde{O}(ε^{-1})$ up bounds in tabular and general function approximation settings, both the concepts of \emph{single-policy concentrability} を与える。
特に、凸解析パイプラインは、ペシミズムの原理を新しい方法で活用することによりこれらの設定を統一し、独立性のある平均値定理に基づいて、過去の研究における証明ルーチンを完全にバイパスする。
さらに、統計率の観点から上界の厳密性を示す、レート最適下界も提供する。
我々の下限は、逆KL正則化と同様に、フォワードKL規則化サンプルの複雑さが低規則化状態における非規則化スローレートを回復することを示した。
関連論文リスト
- Unregularized Convergence of Single-Loop, Entropy-Regularized Natural Actor-Critic [3.756550107432323]
エントロピー正則化は自然政策勾配法の安定化と加速に広く用いられている。
我々は1ループのエントロピー規則化自然アクター・クライトを解析する。
非中心的な批評家を訓練することで、トレーニング方針が決定論に近づいたとしても、我々の批判的追跡は安定し続けることができる。
論文 参考訳(メタデータ) (2026-08-20T03:08:33Z) - Near-Optimal Regret for KL-Regularized Multi-Armed Bandits [54.77408659142336]
KL正規化目標に対するオンライン学習の統計的効率について検討する。
我々は、MABsのKL正規化後悔が$$非依存であることを示し、$tilde(sqrtKT)$とスケールする。
論文 参考訳(メタデータ) (2026-03-02T18:17:33Z) - Regularized Online RLHF with Generalized Bilinear Preferences [68.44113000390544]
一般的な嗜好を伴う文脈的オンラインRLHFの問題を考える。
一般化された双線形選好モデルを用いて、低ランクなスキュー対称行列による選好を捉える。
グリーディポリシーの双対ギャップは推定誤差の正方形によって有界であることを示す。
論文 参考訳(メタデータ) (2026-02-26T15:27:53Z) - Optimal Rates in Continual Linear Regression via Increasing Regularization [39.30412893918111]
本研究では,ランダムなタスク順序付けの下での連続線形回帰について検討する。
この設定では、$k$学習後の最悪の損失は、$Omega (1/k)$の低いバウンドを認める。
明示的等方的$ell$正則化と有限ステップ予算による暗黙的正則化という2つのよく使われる正則化スキームを用いる。
論文 参考訳(メタデータ) (2025-06-06T19:51:14Z) - Logarithmic Regret for Online KL-Regularized Reinforcement Learning [51.113248212150964]
KL正規化は、大規模言語モデルにおけるRL微調整の効率向上に重要な役割を果たしている。
経験的優位性にもかかわらず、KL-正則化RLと標準RLの理論的相違はほとんど未探索のままである。
楽観的なKL正規化オンライン文脈帯域幅アルゴリズムを提案し,その後悔の新たな分析法を提案する。
論文 参考訳(メタデータ) (2025-02-11T11:11:05Z) - Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits [49.96531901205305]
我々は$f$-divergence-regularized offline policy learningを分析する。
逆Kullback-Leibler (KL) の発散に対して、単極集中性の下での最初の$tildeO(epsilon-1)$サンプル複雑性を与える。
これらの結果は,$f$-divergence-regularized policy learningの包括的理解に向けて大きな一歩を踏み出したものと考えられる。
論文 参考訳(メタデータ) (2025-02-09T22:14:45Z) - Shifted Composition III: Local Error Framework for KL Divergence [12.93725028754563]
引数の結合は、2つのプロセス間の偏差を境界付ける中心的なツールである。
カップリングの議論をKL(Kulback-Leibler)の発散に適用する。
論文 参考訳(メタデータ) (2024-12-23T21:40:01Z) - Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation [51.248784084461334]
我々はNesterov加速度アンダーホ条件の一般化版に対する新しい収束率を証明した。
本分析により, 従来の研究に比べて, 強い成長定数への依存度を$$$から$sqrt$に下げることができた。
論文 参考訳(メタデータ) (2024-04-03T00:41:19Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。