論文の概要: Agents That Certify Their Own Exploits: Confidence-Scheduled Restricted Responses for Safe Opponent Exploitation
- arxiv url: http://arxiv.org/abs/2607.28520v1
- Date: Thu, 30 Jul 2026 16:57:57 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-31 21:37:00.68157
- Title: Agents That Certify Their Own Exploits: Confidence-Scheduled Restricted Responses for Safe Opponent Exploitation
- Title(参考訳): 自国の爆発を認証するエージェント:安全対策のための信頼性調整された制限された反応
- Authors: Boning Li, Longbo Huang,
- Abstract要約: 我々は,Emphbudget-Constrained confidence-scheduled limited response (CS-RNR)を導入する。
CS-RNRは、安全保証がエージェントが実際に展開する戦略に基づいて計算する証明書である最初の相手探索方式である。
- 参考スコア(独自算出の注目度): 38.06569764716213
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: An agent playing a Nash-equilibrium strategy in a two-player zero-sum imperfect-information game secures the game value but forfeits the additional value offered by a flawed opponent. Diffuse deviations pose a particular challenge: binary release rules may gather too little evidence to act, while a full best response to an incomplete opponent model can be highly exploitable. We introduce \emph{budget-constrained confidence-scheduled restricted responses} (CS-RNR), the first opponent-exploitation method whose safety guarantee is a certificate the agent computes on the strategy it actually deploys, so that every exploit it commits to is one it has audited itself. The method tracks pooled action frequencies with anytime-valid confidence sequences and treats a frequency as exploitable only once its interval separates from an equilibrium reference. The confirmed deviations define a conservative opponent model, which a restricted-response solve turns into candidate counter-strategies over a grid of pin levels. Before deployment, each complete candidate is evaluated by a full-tree best response. The resulting certificate is compared with a user-specified budget and committed atomically with the strategy. Because this check is performed on the played strategy, model quality determines the exploitation achieved while the certificate controls reference-relative expected loss. In Leduc hold'em, CS-RNR obtains $6.2\times$ the steady-state gain of a money-verified binary gate while keeping every deployed strategy within budget. A trajectory mixture using the same estimator reaches $13.6\times$ the budget. Across Leduc, Liar's Dice, and 5-rank Leduc, all $36{,}000$ audited hands satisfy the reported certificate tolerance.
- Abstract(参考訳): 二プレーヤゼロサム不完全情報ゲームにおいてナッシュ均衡戦略を行うエージェントは、遊技価値を確保できるが、欠陥のある相手の付加価値を無視する。
バイナリリリースルールは、行動する証拠が少なすぎる場合があり、不完全な相手モデルに対する完全なベストな反応は、非常に悪用できる。
我々は,エージェントが実際に展開する戦略について,安全保証が計算する証明書である最初の反対探索手法である 'emph{budget-constrained confidence-scheduled restricted response} (CS-RNR) を紹介する。
この方法は、プールされた動作周波数を任意の時間価の信頼シーケンスで追跡し、その間隔が平衡基準から切り離されたときにのみ、その周波数を悪用できるものとして扱う。
確認された偏差は、制限応答が解ける保守的相手モデルを定義し、ピンレベルのグリッド上の候補反ストラテジーへと変換する。
デプロイ前に、各完全な候補は、フルツリーのベストレスポンスによって評価される。
得られた証明書は、ユーザ指定の予算と比較され、戦略とアトミックにコミットされる。
このチェックはプレイ戦略上で実行されるため、モデル品質は、基準相対的な損失を制御する一方で達成されたエクスプロイトを決定する。
CS-RNR は Leduc hold'em において、全ての配備戦略を予算内に維持しつつ、資金を検証されたバイナリゲートの定常的なゲインを 6.2 の時間で獲得する。
同じ推定器を用いた軌道混合は、予算$13.6\timesに達する。
Leduc、Liar's Dice、および5ランクのLeducの360{,}000ドルの監査金は、報告された証明書の許容度を満たす。
関連論文リスト
- SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute [62.3458279176813]
自己検証リファインメント(Self-Verifying Refinement)は、オラクルフリーのマルチターン強化学習フレームワークである。
自己検証を計算制御ポリシとして使用することを学ぶ。
マクロ平均精度は0.563で、平均で2.99回しか推測できない。
論文 参考訳(メタデータ) (2026-07-30T16:20:58Z) - Gaming-Resistant Insurance Contracts for Autonomous AI Agents: Strategy-Proof Toll Mechanism Design [0.0]
Paper Aは、契約によって固定されたデフォルトに対して各副作用を価格設定し、予備予算に対して実行をゲートする、時間一貫性のあるアクチュアリランタイムを定義する。
我々は、自律型AIエージェント保険契約のための5つの攻撃空間を特徴付け、アクチュアリルランタイムがゲームに耐性があることを証明している。
次に,これらの節をPaper Aのランタイム保証と組み合わせて,5つの攻撃空間に対する共同インセンティブ互換性を得る。
論文 参考訳(メタデータ) (2026-06-15T07:31:21Z) - When Should Agent Trust Be Conditional? Characterizing and Attacking Skill-Conditional Reputation in Agent Swarms [4.9833735627186435]
オープンプラットフォームは、異種エージェント間でタスクをルーティングする傾向にある。
標準評価アプローチは、各エージェントを単一のグローバル信頼スコアで要約する。
技能条件信頼度R(i | k)について検討する。
攻撃者が1つのスキルで安価な証拠を持っていて、ターゲットスキルに誰も条件付きルータをハイジャックしていないことを示す。
論文 参考訳(メタデータ) (2026-06-12T07:32:19Z) - Survive or Collapse: The Asymmetric Roles of Data Gating and Reward Grounding in Self-Play RL [76.45061154544568]
セルフプレイ強化学習は、言語モデルを独自の生成タスクで訓練し、人間ラベルなしでプロジェクタとソルバを共進化させる。
最近のシステムでは強い推理効果が報告されているが、崩壊と不安定性は広く観察され、理解されていない。
代わりに、自己プレイの安定性は、提案者生成タスクがトレーニングプールに入るかを判断するデータレベルゲートと、すでに認められたタスクに関するポリシーを更新する報酬信号の2つの異なるレバーによって管理されていると論じる。
論文 参考訳(メタデータ) (2026-05-21T09:19:23Z) - MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents [0.0]
検索強化エージェントに対するメモリ中毒攻撃を,統合評価フレームワークを用いたStackelbergゲームとして定式化する。
ASR-R: 0.25〜1.00$) による攻撃成功度を4倍に向上させる。
私たちの主な貢献は、勾配結合に接地したキャリブレーションに基づく防御であるMEMSADである。
論文 参考訳(メタデータ) (2026-05-05T08:15:41Z) - ZIP-RC: Optimizing Test-Time Compute via Zero-Overhead Joint Reward-Cost Prediction [57.799425838564]
ZIP-RCは、モデルに報酬とコストのゼロオーバーヘッド推論時間予測を持たせる適応推論手法である。
ZIP-RCは、同じまたはより低い平均コストで過半数投票よりも最大12%精度が向上する。
論文 参考訳(メタデータ) (2025-12-01T09:44:31Z) - Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers [90.50039419576807]
RLVR(Reinforcement Learning with Verifiable Rewards)は、人為的なラベル付けを避けるために、自動検証に対するポリシーを訓練する。
認証ハッキングの脆弱性を軽減するため、多くのRLVRシステムはトレーニング中にバイナリ$0,1$の報酬を破棄する。
この選択にはコストがかかる:textitfalse negatives(正しい回答、FNを拒絶)とtextitfalse positives(間違った回答、FPを受け入れる)を導入する。
論文 参考訳(メタデータ) (2025-10-01T13:56:44Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。