論文の概要: Which Model Is Actually Serving You? IRIS: Budgeted Black-Box Auditing of Model Substitution and Routing Dilution in LLM Gateways
- arxiv url: http://arxiv.org/abs/2607.20860v1
- Date: Thu, 23 Jul 2026 02:32:08 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-24 18:26:25.261901
- Title: Which Model Is Actually Serving You? IRIS: Budgeted Black-Box Auditing of Model Substitution and Routing Dilution in LLM Gateways
- Title(参考訳): IRIS: LLMゲートウェイにおけるモデル置換とルーティング希釈に関する予算付きブラックボックス監査
- Authors: Yuewei Zhang, Zhi-Hai Zhang, Hanzhang Qin,
- Abstract要約: 返却されたテキストのみを必要とする監査である$mathrmIRIS$を提示します。
安価なパイロットは指数的なクエリエラーの崩壊に適合し、疑わしいクエリが発行される前にその予算を凍結する。
$mathrmIRIS$ match or beats detection on shared task, and adapt allocationss the matched-budget target-hit rate from 7,3$% to 8,7$%。
- 参考スコア(独自算出の注目度): 6.964275075635626
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Commercial LLM gateways mediate access to hosted models, but the served backend may not match the advertised one: it may substitute a cheaper model on every request or route only a fraction $ε$ of requests to it. Prior black-box auditors often need a privileged signal (log-probabilities, token ranks, or reference samples) or a target-specific probe, fix the query budget in advance, and return a yes/no verdict. We present $\mathrm{IRIS}$, an audit that needs only the returned text: it asks endpoints to generate random numbers or strings, fingerprints the backend, and is the first to combine, in one text-only audit, detection of whole-stream substitution and fractional dilution, attribution of the served backend, routing-fraction ($ε$) estimation, and a query budget it sizes itself. A cheap pilot fits the exponential query-error decay and freezes that budget before any suspect query is issued. On an intra-family Qwen3 ladder $\mathrm{IRIS}$ verifies the backend at $0.99$ AUROC and sharpens attribution as queries accumulate; across a commercial OpenRouter library it catches $ε{=}0.3$ dilution on margin-qualified pairs at $0.85$ mean power ($0.017$ false-positive rate) and recovers $ε$ to within $0.04$ for enrolled diluents; and a live cross-provider audit flags $14$ of $15$ same-model provider pairs by genuine quantization and kernel deviations, corroborated on third-party MET traces. Against comparable black-box auditors, $\mathrm{IRIS}$ matches or beats detection on shared tasks, and adaptive allocation lifts the matched-budget target-hit rate from $73$% to $87$%. Further experiments cover adversarial gateways, knob identifiability, unseen diluents, and false-positive control.
- Abstract(参考訳): 商用のLLMゲートウェイはホストされたモデルへのアクセスを仲介するが、サービスされたバックエンドは宣伝されたモデルと一致しないかもしれない。
以前のブラックボックス監査人は、しばしば特権的な信号(ログ確率、トークンランク、参照サンプル)やターゲット固有のプローブを必要とし、クエリ予算を事前に修正し、イエス/ノーの判定を返す。
エンドポイントにランダムな数や文字列、指紋、バックエンドを生成するように要求し、1つのテキストのみの監査、ストリーム全体の置換と分数希釈の検出、サービスされたバックエンドの帰属、ルーティング-フレクション(ε$)推定、そしてそれ自身をサイズするクエリ予算を結合する。
安価なパイロットは指数的なクエリエラーの崩壊に適合し、疑わしいクエリが発行される前にその予算を凍結する。
Qwen3 ladder $\mathrm{IRIS}$はバックエンドを$0.99$ AUROCで検証し、クエリの蓄積によって属性を絞り、商用のOpenRouterライブラリでは$ε{=}0.3$の希釈を$0.85$の平均電力$0.017$の偽陽性率で取得し、登録された希釈物に対して$0.04$以内まで回収する。
同等のブラックボックス監査者に対して、$\mathrm{IRIS}$ match or beats detection on shared task、およびアダプティブアロケーションは、マッチした予算のターゲットヒットレートを7,3$%から8,7$%に引き上げる。
さらなる実験では、対向ゲート、ノブ識別性、目に見えない希釈剤、偽陽性制御をカバーしている。
関連論文リスト
- The Security Budget of Code-LLM Prompt Hardening: Provable Limits Under Pass-Only Acceptance [0.0]
本稿では,emphTri-Audit Protocolとしてフロアを運用する。このプロトコルは,プロンプト側推論レジストリ属性をモデル側実証ログから分離する2軸レポーティングプロトコルである。
CodeLlama-7B, Qwen2.5-Coder-7B/1.5B and DeepSeek-Coder-6.7B at $n=164$ yields the emphCross-Model Tri-Audit Invariance: of 28 pass-serving rows, 12-changed-of-record learned-can
論文 参考訳(メタデータ) (2026-06-02T08:22:14Z) - Scaling Laws for Agent Harnesses via Effective Feedback Compute [53.68149869349268]
emphEffective Feedback Compute (EFC)は、情報的、有効、非冗長な場合にのみフィードバックを信用し、その後の決定のために保持するトレースレベルのスケーリング座標である。
EFCベースの座標は、生の計算ベースラインよりも失敗率を常に予測する。
論文 参考訳(メタデータ) (2026-05-28T09:45:47Z) - Auditing Privacy in Multi-Tenant RAG under Account Collusion [1.253312107729806]
マルチテナント検索拡張生成サービスは、ログ単位の差分プライバシーを操作リーク境界として宣伝する。
我々は、同一インデックスのマルチアカウント共謀をプライバシ境界障害とみなす。
修正されていないRAGデプロイメントに対して動作する最初の監査プロトコルを設計する。
論文 参考訳(メタデータ) (2026-05-19T13:41:59Z) - MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents [0.0]
検索強化エージェントに対するメモリ中毒攻撃を,統合評価フレームワークを用いたStackelbergゲームとして定式化する。
ASR-R: 0.25〜1.00$) による攻撃成功度を4倍に向上させる。
私たちの主な貢献は、勾配結合に接地したキャリブレーションに基づく防御であるMEMSADである。
論文 参考訳(メタデータ) (2026-05-05T08:15:41Z) - Correction and Corruption: A Two-Rate View of Error Flow in LLM Protocols [51.56484100374058]
そこで本研究では,単一プロトコルステップを正確なマッチングタスクで監査するためのペアアウトカム計測インタフェースを提案する。
各インスタンスについて、インターフェースはベースラインの正当性ビットと後ステップの正当性ビットを記録する。
これらのレートは精度の変化を予測し、種、混合物、パイプライン間でテスト可能な再利用可能な経験的インターフェースを定義する。
論文 参考訳(メタデータ) (2026-04-20T13:25:40Z) - Consistency-Guided Decoding with Proof-Driven Disambiguation for Three-Way Logical Question Answering [4.878198984786631]
3方向の論理的質問応答(QA)は、$True/False/Unknown$を前提セット$S$の仮説に割り当てる。
CGD-PDは1つの3ウェイ分類器を$H$と$H$の機械式の両方でクエリする軽量なテスト時間層である。
論文 参考訳(メタデータ) (2026-03-12T18:26:16Z) - $V_1$: Unifying Generation and Self-Verification for Parallel Reasoners [69.66089681814013]
$V_$は、効率的なペアワイドランキングを通じて生成と検証を統合するフレームワークである。
V_$-Inferはポイントワイド検証でPass@1を最大10%改善する。
V_$-PairRLは、標準のRLとポイントワイドのジョイントトレーニングよりも、テストタイムのスケーリングが7ドル--9%で向上する。
論文 参考訳(メタデータ) (2026-03-04T17:22:16Z) - Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers [90.50039419576807]
RLVR(Reinforcement Learning with Verifiable Rewards)は、人為的なラベル付けを避けるために、自動検証に対するポリシーを訓練する。
認証ハッキングの脆弱性を軽減するため、多くのRLVRシステムはトレーニング中にバイナリ$0,1$の報酬を破棄する。
この選択にはコストがかかる:textitfalse negatives(正しい回答、FNを拒絶)とtextitfalse positives(間違った回答、FPを受け入れる)を導入する。
論文 参考訳(メタデータ) (2025-10-01T13:56:44Z) - Certifiably Robust Model Evaluation in Federated Learning under Meta-Distributional Shifts [8.700087812420687]
異なるネットワーク "B" 上でモデルの性能を保証する。
我々は、原則付きバニラDKWバウンダリが、同じ(ソース)ネットワーク内の未確認クライアント上で、モデルの真のパフォーマンスの認証を可能にする方法を示す。
論文 参考訳(メタデータ) (2024-10-26T18:45:15Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。