論文の概要: HERALD: Counterfactual Audits and Minimal Repairs for Proof-of-Retrieval Rewards
- arxiv url: http://arxiv.org/abs/2608.06012v1
- Date: Thu, 06 Aug 2026 13:15:54 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-07 15:25:20.912339
- Title: HERALD: Counterfactual Audits and Minimal Repairs for Proof-of-Retrieval Rewards
- Title(参考訳): HERALD:再帰者に対する対人的監査と最小修理
- Authors: Zhuowen Liu, Bohan Cui, YinShang Guo, Yuting Wang, Hao Li,
- Abstract要約: HERALDは、全く同じ調査介入を適用するオフライン監査である。
R[L]$はHotpotQAと2WikiのEM非偽りのゲートと出会うが、MuSiQueではない。
- 参考スコア(独自算出の注目度): 7.036896099993471
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Search-agent rewards mix answer quality, citation grounding, tool cost, and anti-hacking terms; a high score therefore need not imply that cited evidence was retrieved, and added penalties can cancel. We introduce HERALD, an offline audit that applies exact same-question interventions, separates candidate-visible from oracle information, and enumerates detector contracts before policy optimization. On four Qwen3-8B pools from HotpotQA, 2WikiMultiHopQA, and MuSiQue, $R_0$ rejects search deletion and fake IDs, but a label-free citation-laundering attack succeeds. A complete $2^3$ ablation identifies targeted strengthening of $L$---citing a corpus passage absent from the retrieved evidence---as the observed inclusion-minimal repair: $R[L]$ has zero empirical ASR with a 0.50% one-sided cluster upper bound. The gap persists across pool rules, a visible BM25 attacker, and four models; broader hardening remains vulnerable when the attack removes an oracle support-ID penalty. Under strict 5M-token matched training evaluated on 256 paired questions per benchmark, $R[L]$ meets the EM non-inferiority gate on HotpotQA and 2Wiki but not MuSiQue. Equal-suite citation precision and support recall improve by 2.02 and 1.46 points, unsupported citations fall by 1.69, and laundering attackability falls on 2Wiki and MuSiQue. Natural $L$ is not reduced, and the detector appears in only 18 of 58,368 training trajectories. HERALD thus separates robust scoring, sparse learning signal, and policy transfer.
- Abstract(参考訳): 検索エージェントの報酬は、回答の品質、引用の根拠、ツールコスト、アンチハックの用語を混在させ、したがって、引用された証拠が回収されたことを示唆する高いスコアは必要とせず、追加の罰則はキャンセルされる。
HERALDは、全く同一の照会介入を適用し、候補視認性をオラクル情報から切り離し、政策最適化の前に検出器契約を列挙するオフライン監査である。
HotpotQA、2WikiMultiHopQA、MuSiQueの4つのQwen3-8Bプールでは、$R_0$は検索削除と偽のIDを拒否するが、ラベルなしの引用洗浄攻撃は成功する。
完全な2^3$アブレーションは、検出された証拠から欠落したコーパスをアクティベートする$L$の目標強化を識別する。
このギャップは、プールルール、目に見えるBM25攻撃者、および4つのモデルにまたがって持続する。
ベンチマーク毎に256対の質問で評価された厳密な5M-tokenマッチングトレーニングの下では、$R[L]$はHotpotQAと2WikiのEM非不正ゲートに適合するが、MuSiQueではない。
平等な引用精度とサポートリコールは2.02と1.46ポイント改善され、サポートなしの引用は1.69ポイント減少し、洗浄攻撃性は2WikiとMuSiQueに低下する。
天然の$L$は減少せず、検出器は58,368の訓練軌道のうち18にしか現れない。
HERALDは、ロバストスコアリング、スパースラーニング信号、ポリシー転送を分離する。
関連論文リスト
- A Failure-Mode Benchmark for Polymorphic Sybil Poisoning in RAG [4.879573049082894]
そこで我々は, 接地型QAの評価フレームワークと評価フレームワークを, コーディネート検索中毒下でリリースした。
読み手出力を4つの排他的カテゴリ(emphgold, emphhijack, emphabstention, emphdrift)に分割するフレームワーク
本報告では,攻撃目標と攻撃目標を協調的に支援する攻撃クラスであるEmphpolymorphic sybil poisoningを導入する。
論文 参考訳(メタデータ) (2026-07-04T06:56:06Z) - Isolating LLM Alignment from Regex: Zero Coverage and Metric-Dependent Divergence Under Adversarial Mutation [51.56484100374058]
以前の作業では、アクティブフィルタの背後にライブのGeminiバックエンドを追加することで、測定可能なカバレッジが得られなかった。
L_4$-real(Gemini-2.5-flash, token-budget cap, rate limit, output scrub)と同様の$L_5$-no-regexを導入するが、9パターンフィルタは無効である。
3つのサブ言語にまたがる敵対的プローブに対して評価を行った。
論文 参考訳(メタデータ) (2026-06-12T18:54:24Z) - Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing [51.56484100374058]
プロダクションLLMアプリケーションは、いくつかの防衛ファミリを積み重ねる -- 拒絶フレーズフィルタ、トークンバッジコントロール、モデル許容度リスト、レート制限、ツール登録認証 -- が、BASベンチマークでは、単一の集計カバレッジ番号を報告している。
21エージェントベースラインスキャナに4つのLLM-Top-10対応エージェントを追加し、4つの合成LDMエンドポイントの格子をターゲットとした。
論文 参考訳(メタデータ) (2026-06-01T19:39:25Z) - When Do LLM Agents Treat Surface Noise Differently from Semantic Noise? A 68-Cell Measurement Study with a Held-Out Trace-Level Validation [9.055086193088083]
10大言語モデルによって駆動されるチェーン・オブ・シンクとReActエージェントに経験的現象を記述した。
平均的な摂動は、同等の厳しさのプレゼンテーション摂動よりも、最終的な答えを頻繁に変更する。
論文 参考訳(メタデータ) (2026-05-25T15:57:11Z) - MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents [0.0]
検索強化エージェントに対するメモリ中毒攻撃を,統合評価フレームワークを用いたStackelbergゲームとして定式化する。
ASR-R: 0.25〜1.00$) による攻撃成功度を4倍に向上させる。
私たちの主な貢献は、勾配結合に接地したキャリブレーションに基づく防御であるMEMSADである。
論文 参考訳(メタデータ) (2026-05-05T08:15:41Z) - Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers [90.50039419576807]
RLVR(Reinforcement Learning with Verifiable Rewards)は、人為的なラベル付けを避けるために、自動検証に対するポリシーを訓練する。
認証ハッキングの脆弱性を軽減するため、多くのRLVRシステムはトレーニング中にバイナリ$0,1$の報酬を破棄する。
この選択にはコストがかかる:textitfalse negatives(正しい回答、FNを拒絶)とtextitfalse positives(間違った回答、FPを受け入れる)を導入する。
論文 参考訳(メタデータ) (2025-10-01T13:56:44Z) - Eigen-1: Adaptive Multi-Agent Refinement with Monitor-Based RAG for Scientific Reasoning [53.45095336430027]
暗黙的な検索と構造化された協調を組み合わせた統合フレームワークを開発する。
Humanity's Last Exam (HLE) Bio/Chem Goldでは,48.3%の精度を実現している。
SuperGPQAとTRQAの結果はドメイン間の堅牢性を確認した。
論文 参考訳(メタデータ) (2025-09-25T14:05:55Z) - WR-ONE2SET: Towards Well-Calibrated Keyphrase Generation [57.11538133231843]
キーワード生成は、入力文書を要約する短いフレーズを自動的に生成することを目的としている。
最近登場したONE2SETパラダイムは、キーフレーズをセットとして生成し、競争性能を達成した。
本稿では, ONE2SET を拡張した WR-ONE2SET を提案する。
論文 参考訳(メタデータ) (2022-11-13T09:56:24Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。