論文の概要: Corruption-Robust Sparse Linear Contextual Bandits with Knapsack Constraints
- arxiv url: http://arxiv.org/abs/2609.37189v1
- Date: Tue, 29 Sep 2026 10:10:53 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-30 21:28:47.410443
- Title: Corruption-Robust Sparse Linear Contextual Bandits with Knapsack Constraints
- Title(参考訳): Knapsack Constraints による破壊・破壊スパース線形帯域
- Abstract要約: 共同報酬および消費汚職下でのクナプサック制約による疎線形文脈包帯について検討した。
汚職を意識した信頼性幅とオンラインリソース価格と予算安全ルールを組み合わせた推定器・モジュラー・フレームワークを開発した。
- 参考スコア(独自算出の注目度): 6.674723874211299
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: We study sparse linear contextual bandits with knapsack constraints under joint reward and consumption corruption. Consumption corruption creates a challenge beyond corrupted rewards: it affects not only statistical estimates, but also the recorded budget, resource prices, and stopping decisions that govern future allocation. We develop Robust Optimistic Primal--Dual (ROPD), an estimator-modular framework that combines corruption-aware confidence widths with online resource prices and a budget-safety rule. With concrete sparse implementation, ROPD achieves regret against a clean population-LP benchmark of $\widetilde O(T^{2/3}+ΓT^{1/3})$ under forced exploration and population-design coverage, and $\widetilde O(\sqrt T+Γ)$ under on-policy realized-design coverage, for a supplied valid corruption bound $Γ$ under the stated proportional-budget scaling and fixed model/design parameters. When the corruption level is unknown, Shared-Grid adapts confidence radii around common point estimates fitted to a single realized history, incurring explicit initialization and master-comparison costs; its sharper on-policy guarantee additionally requires recommendation coverage. Both methods preserve observed budgets on every realization and bound clean resource violation by cumulative consumption corruption. These results connect corruption-robust sparse estimation with resource accounting, pricing, and stopping in high-dimensional online allocation.
- Abstract(参考訳): 共同報酬および消費汚職下でのクナプサック制約による疎線形文脈包帯について検討した。
消費汚職は、統計的な見積もりだけでなく、記録された予算、資源価格、将来の割り当てを管理する決定の停止にも影響を及ぼす。
汚職を考慮した信頼性幅とオンラインリソース価格と予算安全ルールを組み合わせた推定器・モジュラーフレームワークであるRobust Optimistic Primal-Dual(ROPD)を開発した。
具体的なスパース実装により、ROPDは、強制的な調査と人口設計のカバレッジの下で$$\widetilde O(T^{2/3}+\T^{1/3})$ と$\widetilde O(\sqrt T+)$ のクリーンな人口-LPベンチマークに対して後悔する。
汚職レベルが不明な場合、Shared-Gridは単一の実現した歴史に適合した共通点推定値に信頼度を適応させ、明示的な初期化とマスター比較コストを発生させる。
どちらの手法も、累積消費汚職によるあらゆる実現の予算を維持し、クリーンリソースの侵害を制限している。
これらの結果は,資源会計,価格設定,高次元オンラインアロケーションの停止といった汚職・汚職・スパース推定と結びついている。
関連論文リスト
- Communication-Corruption Coupling and Verification in Cooperative Multi-Objective Bandits [2.1772197319352498]
敵の汚職と限定的検証の下で,ベクトル値の報奨を伴う協調的多腕包帯について検討した。
固定環境サイドの予算$$は、$$から$N$までの効果的な汚職レベルに変換できることを示す。
さらに、避けられない追加的な$()$ペナルティや、クリーンな情報なしではサブ線形後悔が不可能な高破壊体制$=(NT)$など、情報理論の限界を確立する。
論文 参考訳(メタデータ) (2026-01-17T06:13:52Z) - Corruption-Robust Offline Reinforcement Learning with General Function
Approximation [60.91257031278004]
一般関数近似を用いたオフライン強化学習(RL)における劣化問題について検討する。
我々のゴールは、崩壊しないマルコフ決定プロセス(MDP)の最適方針に関して、このような腐敗に対して堅牢で、最適でないギャップを最小限に抑える政策を見つけることである。
論文 参考訳(メタデータ) (2023-10-23T04:07:26Z) - Corruption-Robust Algorithms with Uncertainty Weighting for Nonlinear
Contextual Bandits and Markov Decision Processes [59.61248760134937]
本稿では,$tildeO(sqrtT+zeta)$を後悔するアルゴリズムを提案する。
提案アルゴリズムは、最近開発された線形文脈帯域からの不確実性重み付き最小二乗回帰に依存する。
本稿では,提案アルゴリズムをエピソディックなMDP設定に一般化し,まず汚職レベル$zeta$への付加的依存を実現する。
論文 参考訳(メタデータ) (2022-12-12T15:04:56Z) - Hierarchical Adaptive Contextual Bandits for Resource Constraint based
Recommendation [49.69139684065241]
コンテキスト多重武装バンディット(MAB)は、様々な問題において最先端のパフォーマンスを達成する。
本稿では,階層型適応型文脈帯域幅法(HATCH)を提案する。
論文 参考訳(メタデータ) (2020-04-02T17:04:52Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。