論文の概要: When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR
- arxiv url: http://arxiv.org/abs/2607.11022v1
- Date: Mon, 13 Jul 2026 02:41:16 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-14 17:47:21.305949
- Title: When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR
- Title(参考訳): Reward Suiteがリークされた時: RLVRにおける自然検証用偽陽性者の事前登録された因果関係
- Authors: Chuyifei Zhang,
- Abstract要約: コードに対するRLVR報酬として使用されるテストスイートには、自然な偽陽性がある。
デプロイされたスイート上で、事前登録された2本腕の因果コントラストを実行します。
報酬は単なるスイートアーティファクトではなく、実際のバグに対して支払われていることに気付きました。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: The test suites used as RLVR rewards for code have natural false positives: per-task, persistent, asymmetric errors that accept the same wrong programs every time they appear, unlike the symmetric or resampled noise assumed by existing noise-robustness analyses. We run a preregistered two-arm causal contrast on a deployed suite: GRPO on identical MBPP tasks, seeds, and compute, rewarded by the original MBPP tests (leaky) versus the MBPP+ extra tests (hardened). Two further families replicate the design under a preregistration frozen before their data existed. [C] The average held-out effect is bounded: non-inferior under a preregistered 1.5-pt margin (gap 0.20 pt, one-sided 95% upper bound 0.75 pt). [C] Rewarded false-positive mass tracks a cheap static leakiness audit computed before training (Spearman 0.80), and the registered train-side test puts the leak-stratum FP share +43.8 pt above clean tasks. [E] Auditing every rewarded FP under signed, human-adjudicated rules finds a large residual of verified genuinely wrong code: 47.57% record-weighted; both replication families reproduce a large share. The reward paid for real bugs, not merely suite artifacts. [E] Mechanism evidence is consistent with selection of pre-existing error modes rather than learned exploitation: FP incidence does not grow within our horizon, and untrained base models already produce the same wrong outputs under the leaky filter. We then turn the same instrument on the frontier judges themselves: on their own false positives they self-assess only weakly, a same-author test is unresolved, and even the highest-scoring reader we probe stays far below its score on a weaker policy's errors -- two subjects on MBPP, licensing nothing about frontier models in general. A cheap static audit locates exposure before training; hardening the reward removes the measurement inflation, though here it buys little capability.
- Abstract(参考訳): コードに対するRLVR報酬として使用されるテストスイートには自然な偽陽性がある: タスク毎、永続的、非対称的エラーは、出現するたびに同じ間違ったプログラムを受理する。
既存のMBPPテスト(リーキー)とMBPP+の余分なテスト(強化)により報われる、同一のMBPPタスク、種、および計算上のGRPOを、デプロイスイート上で事前登録した2本腕因果コントラストを実行する。
さらに2つの家族が、データが存在する前に凍結した事前登録の下で設計を複製した。
[C]平均ホールドアウト効果は、1.5-ptの差(ギャップ0.20 pt、片側95%の上界0.75 pt)の下で非インフェニオールで有界である。
[C]逆偽陽性マスは、トレーニング前に計算した安価な静的な漏れ度監査(Spearman 0.80)を追跡し、登録された列車側試験は、リーク層FPのシェア+43.8 ptをクリーンタスク上に置く。
Audiing every rewarded FP under signed, human-adjudicated rules found a lot of certainly wrong code: 47.57% record-weighted; both replication family repeat a large share。
報酬は、単なるスイートアーティファクトではなく、本物のバグに対して支払われた。
FPの出現は我々の地平線内では増加せず、訓練されていないベースモデルは、漏れやすいフィルタの下で既に同じ間違った出力を生成する。
次に、フロンティアの裁判官自身で同じ楽器を回す。彼ら自身の偽陽性では、自己評価は弱く、同じ著者によるテストは未解決であり、調査する最高評価の読者でさえ、弱い政策のエラーでスコアをはるかに下回っている -- MBPPの2人の被験者は、フロンティアのモデル全般について何もライセンスしていない。
安価な静的監査はトレーニング前に露光を検知するが、報酬の硬化は測定インフレーションを除去する。
関連論文リスト
- Certified Finite-Shot Operating Windows for Virtual Distillation and Symmetry Verification [0.0]
仮想蒸留 (VD) と対称性検証 (SV) で比較可能な有限ショット動窓理論を開発した。
VD の場合、この法則は商推定器の統計バイアスと分母不安定性を捉え、商が信頼できる範囲を超えてサンプルサイズを特定する濃度証明書を定めている。
SVでは、検出不能なエラーによって残されたバイアスフロアと、受信確率によって設定されたサンプリングペナルティを分離する。
論文 参考訳(メタデータ) (2026-06-13T20:42:40Z) - The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning [54.0056619145783]
プロセス・リワード・モデル (Process Reward Models, PRM) は、段階的なフィードバックを提供することで、推論のためのクレジット割り当てを改善する。
ステップレベルのトレーニングデータにおいて,重度の不均衡に起因するPRMの隠れバイアスを同定する。
対照的な段階比較から学習するポリシ対応のPRMトレーニングフレームワークであるPRISMを提案する。
論文 参考訳(メタデータ) (2026-06-08T06:22:33Z) - Retrying vs Resampling in AI Control [0.42970700836450476]
我々は、AI制御の観点から再試行を行い、モデルが潜在的に敵対的なものとして扱う。
再試行は正直な疑念のスコアを減少させるが、信頼できないモデルは監視の合理性を利用してスニーカー攻撃を構築することができる。
論文 参考訳(メタデータ) (2026-05-25T17:10:41Z) - Correction and Corruption: A Two-Rate View of Error Flow in LLM Protocols [51.56484100374058]
そこで本研究では,単一プロトコルステップを正確なマッチングタスクで監査するためのペアアウトカム計測インタフェースを提案する。
各インスタンスについて、インターフェースはベースラインの正当性ビットと後ステップの正当性ビットを記録する。
これらのレートは精度の変化を予測し、種、混合物、パイプライン間でテスト可能な再利用可能な経験的インターフェースを定義する。
論文 参考訳(メタデータ) (2026-04-20T13:25:40Z) - CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal [84.71254539482369]
検証可能な報酬を伴うグループ相対的強化学習(RLVR)は、しばしば、すでに失敗している最も情報に富むデータを浪費する。
エラーを監督するマルチモーダル推論のための,障害中心のポストトレーニングフレームワークであるCAREを提案する。
CAREは正確さを改善し、スムーズさをトレーニングすると同時に、障害からの学習信号のシェアを明示的に増やします。
論文 参考訳(メタデータ) (2025-12-22T16:34:21Z) - Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers [90.50039419576807]
RLVR(Reinforcement Learning with Verifiable Rewards)は、人為的なラベル付けを避けるために、自動検証に対するポリシーを訓練する。
認証ハッキングの脆弱性を軽減するため、多くのRLVRシステムはトレーニング中にバイナリ$0,1$の報酬を破棄する。
この選択にはコストがかかる:textitfalse negatives(正しい回答、FNを拒絶)とtextitfalse positives(間違った回答、FPを受け入れる)を導入する。
論文 参考訳(メタデータ) (2025-10-01T13:56:44Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。