論文の概要: Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades
- arxiv url: http://arxiv.org/abs/2609.01345v1
- Date: Tue, 01 Sep 2026 14:53:41 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-02 16:31:36.781985
- Title: Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades
- Title(参考訳): チープ検証器と大型盲点:コストセービングカスケードの信頼性コストの測定
- Authors: Dushyant Rajput,
- Abstract要約: 推論カスケードは、ほとんどのクエリを安価なモデルで答え、検証として機能するフロンティアモデルにハードテールをエスカレートすることでコストを削減する。
検証者の拒否に対して安価な学生を微調整することで、エスカレーション率とコストは各ラウンドで低下する。
このループを実LLMで測定し,4つの知見を報告する。
- 参考スコア(独自算出の注目度): 0.0
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Inference cascades cut cost by answering most queries with a cheap model and escalating a hard tail to a frontier model that acts as verifier. A natural extension closes the loop: fine-tune the cheap student on the verifier's rejections so the escalation rate, and cost, fall each round. We measure this loop on real LLMs and report four findings. First, the verifier's blind spot, the fraction of the student's wrong answers it accepts, is large and moves adversarially: it grows with student capability ($β$ from 0.12 to 0.55 as the student scales 0.5B to 32B) and shrinks with verifier capability, so it is worst in the cheap-student, cheap-verifier regime cascades exist to create. Second, buying it away returns the saving: a frontier verifier drives $β$ to about 0.05 but then escalates on 46% of hard-MATH queries against a 39% true error rate, paying the frontier price on nearly half of all traffic. Third, naive corrective fine-tuning on the verifier-rejected tail does not improve the small student but degrades and ultimately collapses it, across every teacher we tried (cross-family and same-family), so at this scale the self-improving loop is self-defeating. Fourth, through all of this the cascade's own dashboard, every metric computed through the verifier, reads a flat 3% error while true delivered error swings up to 32%: the system is blind to its own degradation by construction. We then give the theory that explains the blindness, a two-population conservation law, $ε_\infty \lesssim q_0 β_0$, under which every in-loop metric improves while true quality does not, and a synthetic study that validates the mechanism. The practical conclusion: the reliability of a self-improving cascade cannot be read from any metric computed through its own verifier.
- Abstract(参考訳): 推論カスケードは、ほとんどのクエリを安価なモデルで答え、検証として機能するフロンティアモデルにハードテールをエスカレートすることでコストを削減する。
検証者の拒否に対して安価な学生を微調整することで、エスカレーション率とコストは各ラウンドで低下する。
このループを実LLMで測定し,4つの知見を報告する。
学生の能力(0.5Bから32Bにスケールするとβ$0.12から0.55)で成長し、検証能力で縮小するので、安価で安価な検証体制のカスケードでは最悪のものとなる。
フェデラル検証器は$β$を約0.05に駆動するが、ハードMATHクエリの46%で39%の真のエラー率に対してエスカレートし、全トラフィックのほぼ半分でフデラー価格を支払う。
第3に,検証対象の尾部における直感的な修正的微調整は,小学生を向上させるものではないが,最終的に劣化し,崩壊させる。
第4に、カスケード自身のダッシュボードは、バリデーションを通じて計算されたすべてのメトリックで、フラット3%のエラーを読み、真のデリバリエラーは最大32%までスイングする。
次に、盲性を説明する理論と、2つの人口保存法、$ε_\infty \lesssim q_0 β_0$ を与える。
現実的な結論:自己改善カスケードの信頼性は、自分自身の検証器によって計算された任意の計量から読めない。
関連論文リスト
- The Constitutional Coverage Trilemma in AI Governance [0.0]
AIシステムは、Emphconstitutional institutionsとして機能する。各デプロイされたモデルは、安全性、有用性、誠実性、自律性、株式の暗黙のランクを符号化する。
我々は、フロンティア憲法型の供給が人間の需要をカバーするかどうかを問う。
論文 参考訳(メタデータ) (2026-09-01T14:08:56Z) - Teaching a Large Language Model Tutor to Withhold the Answer: A Supervisor Architecture and an Evidence-Driven Method for Tuning Socratic Behavior [0.0]
効果的な大規模言語モデル (LLM) チューターは、容易に作成できる答えを与えるために、しばしば辞退しなければならない。
本稿では,ターン毎の機械チェック可能な契約として回答保持を強制する教育システムについて報告する。
人間の被験者を使わない自動評価で行動を調整する。
論文 参考訳(メタデータ) (2026-08-12T17:35:58Z) - What Actually Works for Spacecraft Fault-Tolerant Control: An Honest Settled-Gate Benchmark of Learned and Classical Methods [2.28438857884398]
近年のFTCの研究は、宇宙船のアクチュエーターの故障で高い成功を収めたことを報告している。
我々は、成功がトレーニングで見たことのない欠陥にそれを保持することを意味しているとき、宇宙船が何を指しているのかを尋ねる。
私たちは、落ち着いたゲートの周りに構築されたベンチマークで回答します。
論文 参考訳(メタデータ) (2026-06-24T04:08:39Z) - Efficiently Learning Drifting Halfspaces with Massart Noise [50.4331323695175]
本研究では,マッサートノイズの存在下での漂流概念の学習問題について検討する。
このフレームワークでは、オンライン学習者は独立したサンプルの履歴にアクセスすることができる。
目標は、各ラウンドで小さな予測誤差の仮説を出力することである。
論文 参考訳(メタデータ) (2026-06-09T17:35:18Z) - Teacher-Free Self-Training Amplifies but Does Not Compound: A Pass@$K$ Crossover on a Free-Verifier Domain [0.0]
教師なしの「コンステレーション」でその疑問を解き明かす
すべてが4ビットのQwen3-4Bと24GBのGPUで動いています。
論文 参考訳(メタデータ) (2026-06-05T21:37:49Z) - Causal Label Recovery in Payment Networks [0.0]
支払いネットワークにおける不正検出モデルは、体系的にバイアスのあるチャージバックラベルでトレーニングする。
共用紙 [arXiv05:26.27557] は、これらの4つの障害が検出性能に最小限の上限を課すことを示した。
観測パイプラインを3段階, 破損層を有する逐次欠落データ問題として定式化し, 逐次トリプライロバスト推定器を構築した。
論文 参考訳(メタデータ) (2026-05-28T02:43:34Z) - From Patches to Trajectories: Privileged Process Supervision for Software-Engineering Agents [56.31499185764872]
教師の長い軌道上の監督された微調整(SFT)は、オープンソフトウェアエンジニアリング(SWE)エージェントに調査と推論を浸透させる主要な方法である。
本稿では,P2T (Patches-to-Trajectories) を提案する。P2T (Patches-to-Trajectories) は,P2T (Patches-to-Trajectories) において,P2T (Patches-to-Trajectories) とP2T (Patches-to-Trajectories) の2つの最適化法である。
論文 参考訳(メタデータ) (2026-05-21T04:54:55Z) - CoVerRL: Breaking the Consensus Trap in Label-Free Reasoning via Generator-Verifier Co-Evolution [52.691495954442985]
CoVerRLは1つのモデルがジェネレータと検証ロールを交換するフレームワークで、各機能が他方をブートストラップする。
Qwen と Llama のモデルファミリーでの実験では、CoVerRL は数理推論のベンチマークで4.7-5.9% でラベルなしのベースラインを上回っている。
自己検証の精度は55%から85%以上改善され、両方の能力が真に共存することを確認した。
論文 参考訳(メタデータ) (2026-03-18T14:38:55Z) - GATES: Self-Distillation under Privileged Context with Consensus Gating [89.62339954332248]
我々は、監督が信頼できない環境で自己蒸留を研究する。
非対称な文脈で回答する文書に焦点をあてる。
複数の文書ベース推論トレースをサンプリングすることにより、教師のコンセンサスからオンラインでの監督を導出する。
論文 参考訳(メタデータ) (2026-02-24T05:56:20Z) - Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers [90.50039419576807]
RLVR(Reinforcement Learning with Verifiable Rewards)は、人為的なラベル付けを避けるために、自動検証に対するポリシーを訓練する。
認証ハッキングの脆弱性を軽減するため、多くのRLVRシステムはトレーニング中にバイナリ$0,1$の報酬を破棄する。
この選択にはコストがかかる:textitfalse negatives(正しい回答、FNを拒絶)とtextitfalse positives(間違った回答、FPを受け入れる)を導入する。
論文 参考訳(メタデータ) (2025-10-01T13:56:44Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。