論文の概要: Statistical Efficiency and Inference of Quantile Distributional Reinforcement Learning
- arxiv url: http://arxiv.org/abs/2607.08444v1
- Date: Thu, 09 Jul 2026 13:06:07 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-10 14:45:27.543597
- Title: Statistical Efficiency and Inference of Quantile Distributional Reinforcement Learning
- Title(参考訳): 量子分布強化学習の統計的効率と推論
- Abstract要約: 統計的効率の観点から, 量子化に基づく分布強化学習について検討する。
W_infty$ の上限の下で $_m(n)$ と $_m$ の非漸近誤差を定めている。
量子化に基づく推定器は、無限次元の極限においてパラメトリックに効率的であることを示す。
- 参考スコア(独自算出の注目度): 16.700969464418936
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: In this paper, we study quantile-based distributional reinforcement learning from the perspective of statistical efficiency. We focus on distributional policy evaluation, whose goal is to characterize the return distribution, namely the distribution of discounted cumulative rewards under a given policy. To obtain a finite-dimensional representation of the return distribution, we consider the quantile fixed point $η_m$ induced by the quantile-projected distributional Bellman equation. Assuming access to a generative model, we construct an estimator $η_m^{(n)}$ based on an empirical Markov decision process. For a fixed number of quantiles $m$, we establish a non-asymptotic error bound for $η_m^{(n)}$ and $η_m$ under the supremum $W_\infty$ metric, showing that the estimation error scales as $\widetilde{O}(\sqrt{m/n})$ with respect to $m$ and $n$. This implies that the quantile-based distributional policy evaluation problem can be solved with sample efficiency, achieving the optimal parametric $\sqrt{n}$ convergence rate. We derive the asymptotic distribution of the quantile parameters $\sqrt{n}(θ_m^{(n)}-θ_m)$ and characterize the semiparametric efficiency bound, which is attained by our estimator. Beyond the fixed-dimensional setting, we investigate the asymptotic regime in which the number of quantiles diverges. We characterize the limit covariance structure and show that it matches the semiparametric efficiency bound of the nonparametric model for distributional policy evaluation, showing that quantile-based estimators remain asymptotically efficient in the infinite-dimensional limit. Finally, we establish a Berry--Esseen theorem for smooth functionals $\sqrt{n}(η_m^{(n)}(s)-η_m(s))f$, thereby providing a foundation for statistically valid inference on functionals of the quantile-projected return distribution.
- Abstract(参考訳): 本稿では,統計的効率の観点から,量子量に基づく分散強化学習について検討する。
我々は,返却分布,すなわち所定の方針の下での累積報酬の分配を特徴付けることを目標とする配当政策評価に焦点をあてる。
帰納分布の有限次元表現を得るために、量子射影分布ベルマン方程式によって誘導される量子的不動点 $η_m$ を考える。
生成モデルへのアクセスを仮定すると、経験的マルコフ決定プロセスに基づいて推定子$η_m^{(n)}$を構築する。
固定数の量子化に対して、$m$に対して$η_m^{(n)}$と$η_m$の非漸近誤差を確立し、$m$と$n$に対して$\widetilde{O}(\sqrt{m/n})$として推定誤差がスケールすることを示す。
これは、Quantile-based distributional policy evaluation problemをサンプル効率で解くことができ、最適なパラメトリック $\sqrt{n}$ Concurence rate を達成することを意味する。
我々は、量子化パラメータ $\sqrt{n}(θ_m^{(n)}-θ_m)$ の漸近分布を導出し、その半パラメトリックな効率境界を特徴づける。
固定次元設定の他に、量子の数が分岐する漸近的状態について検討する。
限界共分散構造を特徴付けるとともに,非パラメトリックモデルの半パラメトリック効率境界と分布ポリシ評価を一致させ,無限次元極限において量子的推定器が漸近的に効率を保っていることを示す。
最後に、滑らかな函数 $\sqrt{n}(η_m^{(n)}(s)-η_m(s))f$ に対するベリー-エッシーの定理を確立し、したがって、量子射影された戻り分布の関数に対する統計的に有効な推論の基礎を与える。
関連論文リスト
- On Stability and Decomposition of Sample Quantiles under Heavy-Tailed Distributions [0.0]
経験過程理論は、半空間の力学、対称差分、およびGlivenko-Cantelli一様収束を通した使用可能な足場を提供する。
本稿では、投影方向と量子閾値効果を分離するQ-Qityの定式化を提案する。
論文 参考訳(メタデータ) (2026-05-18T13:19:29Z) - Non-asymptotic quantisation of spherically symmetric distributions [1.2031796234206138]
ザドールの有名な定理は最適な量子化の基礎である。
適半径$r$の球面上に均一に分布する適度な$n$ランダム量子化器では、例外的な性能が得られることを示す。
特に$n$が$d$でスケールするシナリオでは、$r$の近似を導出します。
論文 参考訳(メタデータ) (2026-05-12T10:01:41Z) - Stabilizing Fixed-Point Iteration for Markov Chain Poisson Equations [49.702772230127465]
有限状態マルコフ鎖を$n$状態と遷移行列$P$で研究する。
すべての非退化モードが実周辺不変部分空間 $mathcalK(P)$ によってキャプチャされ、商空間 $mathbbRn/mathcalK(P) 上の誘導作用素が厳密に収縮し、ユニークな商解が得られることを示す。
論文 参考訳(メタデータ) (2026-01-31T02:57:01Z) - Approximating $f$-Divergences with Rank Statistics [0.3222802562733787]
ランクの分布を直接扱うことで、明示的な密度比推定を避けるために、$f$-divergencesのランク統計近似を導入する。
発散の結果として生じる推定量は、K$の単調であり、常に真$f$-発散の下位境界であることを示す。
ニューラルベースラインに対するベンチマークによるアプローチを実証的に検証し,生成モデル実験における学習目的としての利用を例証する。
論文 参考訳(メタデータ) (2026-01-30T10:05:33Z) - Beyond likelihood ratio bias: Nested multi-time-scale stochastic approximation for likelihood-free parameter estimation [49.78792404811239]
確率分析形式が不明なシミュレーションベースモデルにおける推論について検討する。
我々は、スコアを同時に追跡し、パラメータ更新を駆動する比率のないネスト型マルチタイムスケール近似(SA)手法を用いる。
我々のアルゴリズムは、オリジナルのバイアス$Obig(sqrtfrac1Nbig)$を排除し、収束率を$Obig(beta_k+sqrtfracalpha_kNbig)$から加速できることを示す。
論文 参考訳(メタデータ) (2024-11-20T02:46:15Z) - Kernel-based off-policy estimation without overlap: Instance optimality
beyond semiparametric efficiency [53.90687548731265]
本研究では,観測データに基づいて線形関数を推定するための最適手順について検討する。
任意の凸および対称函数クラス $mathcalF$ に対して、平均二乗誤差で有界な非漸近局所ミニマックスを導出する。
論文 参考訳(メタデータ) (2023-01-16T02:57:37Z) - Optimal policy evaluation using kernel-based temporal difference methods [78.83926562536791]
カーネルヒルベルト空間を用いて、無限水平割引マルコフ報酬過程の値関数を推定する。
我々は、関連するカーネル演算子の固有値に明示的に依存した誤差の非漸近上界を導出する。
MRP のサブクラスに対する minimax の下位境界を証明する。
論文 参考訳(メタデータ) (2021-09-24T14:48:20Z) - Heavy-tailed Streaming Statistical Estimation [58.70341336199497]
ストリーミング$p$のサンプルから重み付き統計推定の課題を考察する。
そこで我々は,傾きの雑音に対して,よりニュアンスな条件下での傾きの傾きの低下を設計し,より詳細な解析を行う。
論文 参考訳(メタデータ) (2021-08-25T21:30:27Z) - Understanding the Under-Coverage Bias in Uncertainty Estimation [58.03725169462616]
量子レグレッションは、現実の望ましいカバレッジレベルよりもアンファンダーカバー(enmphunder-cover)する傾向がある。
我々は、量子レグレッションが固有のアンダーカバーバイアスに悩まされていることを証明している。
我々の理論は、この過大被覆バイアスが特定の高次元パラメータ推定誤差に起因することを明らかにしている。
論文 参考訳(メタデータ) (2021-06-10T06:11:55Z) - Estimation in Tensor Ising Models [5.161531917413708]
N$ノード上の分布から1つのサンプルを与えられた$p$-tensor Isingモデルの自然パラメータを推定する問題を考える。
特に、$sqrt N$-consistency of the MPL estimate in the $p$-spin Sherrington-Kirkpatrick (SK) model。
我々は、$p$-tensor Curie-Weiss モデルの特別な場合における MPL 推定の正確なゆらぎを導出する。
論文 参考訳(メタデータ) (2020-08-29T00:06:58Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。