論文の概要: Vector Bellman Theory for Multichain Robust Average-Reward Markov Decision Processes
- arxiv url: http://arxiv.org/abs/2609.28792v1
- Date: Wed, 23 Sep 2026 21:09:55 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-25 21:10:09.558523
- Title: Vector Bellman Theory for Multichain Robust Average-Reward Markov Decision Processes
- Title(参考訳): 多鎖ロバスト平均逆マルコフ決定過程のベクトルベルマン理論
- Abstract要約: コンパクトで後作用の$(s,a)$-矩形曖昧性を持つ有限モデルに対するベクトルベルマン理論を開発する。
ゲインファースト、バイアス秒最適化の原則は、結合ベクターゲインバイアスシステムをもたらす。
有限ベルマン可解性の下では、利得推定とベルマン変位は最適利得ベクトルに収束する。
- 参考スコア(独自算出の注目度): 2.6610322177713512
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Robust average-reward Markov decision processes provide a fundamental framework for long-term performance optimization under uncertainty, and can have optimal long-run rewards that depend on the initial state. This state dependence requires a vector Bellman theory that accounts for both recurrent-class rewards and transition uncertainty. We develop such a theory for finite models with compact, post-action $(s,a)$-rectangular ambiguity. A gain-first, bias-second optimization principle yields a coupled vector gain-bias system, and every finite solution identifies the optimal robust gain and supplies stationary saddle strategies against history-dependent opponents, simultaneously from all initial states. We further characterize solvability through stationary gain conditions and a uniform bound on canonical transient corrections, and give sufficient conditions that permit distinct recurrent-class gains. The certificates also yield asymptotically affine trajectories of the robust Bellman operator, based on which we design a robust approximately shifted Halpern planning algorithm. Under finite Bellman solvability, the gain estimates and Bellman displacements converge to the optimal gain vector, and every extracted greedy controller is average-optimal after a finite, instance-dependent budget. These results thus connect finite Bellman certificates to undiscounted planning for state-dependent robust average rewards, providing theoretical understandings.
- Abstract(参考訳): ロバストな平均回帰マルコフ決定プロセスは、不確実性の下での長期的なパフォーマンス最適化のための基本的なフレームワークを提供し、初期状態に依存する最適な長期報酬を得ることができる。
この状態依存は、再帰的なクラス報酬と遷移の不確実性の両方を考慮に入れたベクトルベルマン理論を必要とする。
コンパクトで後作用の $(s,a)$-正方形のあいまいさを持つ有限モデルに対するそのような理論を開発する。
ゲインファースト、バイアス秒の最適化原理は結合ベクターゲインバイアスシステムをもたらし、全ての有限解は最適なロバストゲインを特定し、すべての初期状態から同時に履歴に依存した相手に対して定常サドル戦略を供給する。
さらに、定常ゲイン条件と正準過渡補正に一様に拘束された均一な解法を特徴付け、異なる再帰クラスゲインを許容する十分な条件を与える。
証明はまた、ロバストなベルマン作用素の漸近的なアフィン軌道をもたらし、Halpern計画アルゴリズムを設計する。
有限ベルマン可解性の下では、ゲイン推定とベルマン変位は最適ゲインベクトルに収束し、抽出されたグリーディ制御子は、有限のインスタンス依存予算の後に平均最適となる。
これらの結果は、有限ベルマン証明書と、状態依存の頑健な平均報酬の未公表の計画とを結びつけ、理論的理解を提供する。
関連論文リスト
- Central Limit Theorem for Two-Time-Scale Approximate Distributionally Robust RL [8.809468023364703]
堅牢な強化学習アルゴリズムの設計は根本的な課題を生んでいる。
本稿では、関連するロバスト関数の1次展開に基づく近似DRRLフレームワークを提案する。
この近似方程式の定点を学習するために,平均変数近似(MVSA)を提案する。
論文 参考訳(メタデータ) (2026-05-08T19:24:28Z) - Contraction-Aligned Analysis of Soft Bellman Residual Minimization with Weighted Lp-Norm for Markov Decision Problem [9.333190920811626]
ベルマン残差最小化のソフトな定式化を検討し、一般化された重み付きLp-ノルムに拡張する。
p が増加するにつれて、この定式化はベルマン作用素の縮約幾何と最適化目標を一致させることを示す。
本分析は,残差最小化とベルマン縮約の原理的接続を提供し,誤差伝搬の制御を改良する。
論文 参考訳(メタデータ) (2026-04-08T08:58:30Z) - Robust and Consistent Ski Rental with Distributional Advice [13.811651343801579]
スキーレンタル問題は不確実性の下でのオンライン意思決定の標準モデルである。
本稿では,未知品質の分布的アドバイスを決定論的アルゴリズムとランダム化アルゴリズムの両方に統合するフレームワークを提案する。
我々のフレームワークは、既存の点予測ベースラインよりも一貫性を著しく向上し、かつ、同等の堅牢性を維持していることを示す。
論文 参考訳(メタデータ) (2026-03-31T04:04:21Z) - Stopping Rules for Stochastic Gradient Descent via Anytime-Valid Confidence Sequences [51.56484100374058]
凸最適化のための勾配降下(SGD)の停止規則について検討した。
我々は、投影されたSGDの重み付き平均準最適度に対して、任意の有意、データ依存の高信頼シーケンスを開発する。
これらは、厳格でタイムユニフォームなパフォーマンス保証と、有限時間$varepsilon$-optimality証明書である。
論文 参考訳(メタデータ) (2025-12-15T09:26:45Z) - ZIP-RC: Optimizing Test-Time Compute via Zero-Overhead Joint Reward-Cost Prediction [57.799425838564]
ZIP-RCは、モデルに報酬とコストのゼロオーバーヘッド推論時間予測を持たせる適応推論手法である。
ZIP-RCは、同じまたはより低い平均コストで過半数投票よりも最大12%精度が向上する。
論文 参考訳(メタデータ) (2025-12-01T09:44:31Z) - Regularized Q-Learning with Linear Function Approximation [2.765106384328772]
線形汎関数近似を用いた正規化Q-ラーニングの2段階最適化について検討する。
特定の仮定の下では、提案アルゴリズムはマルコフ雑音の存在下で定常点に収束することを示す。
論文 参考訳(メタデータ) (2024-01-26T20:45:40Z) - Bayesian Bellman Operators [55.959376449737405]
ベイズ強化学習(RL)の新しい視点について紹介する。
我々のフレームワークは、ブートストラップが導入されたとき、モデルなしアプローチは実際には値関数ではなくベルマン作用素よりも後部を推測する、という洞察に動機づけられている。
論文 参考訳(メタデータ) (2021-06-09T12:20:46Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。