論文の概要: SIFT: Robust Meta-Faithfulness Verification of Chain-of-Thought Reasoning Under Distribution Shift
- arxiv url: http://arxiv.org/abs/2610.04594v1
- Date: Sat, 03 Oct 2026 15:30:34 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-11 18:35:45.391573
- Title: SIFT: Robust Meta-Faithfulness Verification of Chain-of-Thought Reasoning Under Distribution Shift
- Title(参考訳): SIFT:配電シフト下におけるチェーンオブソート推論のロバストなメタフェースフルネス検証
- Abstract要約: CoT(Chain-of-Thought)忠実度検出器は推論モデルの監査に広く用いられている。
我々は、検出器が分布シフト中、自分自身に忠実であるかどうかを問う。
- 参考スコア(独自算出の注目度): 0.6465251961564605
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Chain-of-Thought (CoT) faithfulness detectors are widely used to audit reasoning models, yet a detector is itself a predictor whose verdicts are treated as stable properties. We ask whether a detector is faithful to itself under distribution shift. We formalize meta-faithfulness as an invariance principle: a valid detector must return identical verdicts on traces that differ only by transformations preserving ground-truth faithfulness. We prove three results: (i) no detector using only intervention-response profiles can separate faithful from epiphenomenal mechanisms with identical signatures; (ii) any detector relying on shift-sensitive features violates invariance at a rate independent of its in-distribution accuracy; (iii) an asymptotic certified selective-risk guarantee enables confident abstention. We operationalize the principle in FaithShift, a stress-test protocol spanning ten shift axes, and propose SIFT, a hidden-state trajectory detector trained with cross-environment invariance objectives and certified abstention. Across 14,996 traces, four domains, and eight models, three findings emerge. First, transfer collapse is real: all existing detectors show gaps $\geq 0.15$ AUROC. Second, the dominant bottleneck is sampling stochasticity, not shift: over 80% of detector instability stems from random seed variation, falsifying our preregistered prediction that shift-attributable violations exceed 0.25. Third, SIFT cuts invariance violations by 64% over the best single-seed baseline, but a four-seed ensemble of any detector narrows the margin to 0.01 (indistinguishable at matched coverage, $p=0.21$), and SIFT needs a 51% abstention rate. Cross-model transfer degrades from within-family to cross-family to open-weight-to-API, partly closed by multi-model training. We offer a framework for auditing auditors: the real barrier is detector variance, not distribution shift.
- Abstract(参考訳): CoT(Chain-of-Thought)忠実度検出器は推論モデルの監査に広く用いられているが、検出器自体が安定な特性として評価される予測器である。
我々は、検出器が分布シフト中、自分自身に忠実であるかどうかを問う。
我々はメタ忠実性を不変原理として定式化する: 有効な検出器は、基底真正性を保存する変換によってのみ異なるトレース上の同一の検証を返さなければならない。
3つの結果を証明します。
i) 介入応答プロファイルのみを用いた検出装置は、同一のシグネチャを有するてんかん機構から忠実に分離することができません。
二 シフト感度特性に依存する検出器は、その分布内精度によらない速度で不均一を犯す。
三 適格選択リスク保証により、確実な棄権が可能であること。
本研究では,10個のシフト軸にまたがるストレステストプロトコルであるFaithShiftの原理を運用し,クロス環境分散目標と認定棄却条件を訓練した隠れ状態軌道検出器SIFTを提案する。
14,996個の痕跡、4つのドメイン、8つのモデルがあり、3つの発見が浮かび上がっている。
まず、転送崩壊は現実的であり、既存の検出器は全てギャップを$\geq 0.15$ AUROCとしている。
第2に、シフトではなく確率性のサンプリングが主なボトルネックであり、検出器不安定性の80%以上がランダムな種の変化によるもので、シフト帰属的違反が0.25を超えるという事前登録された予測を偽る。
第3に、SIFTは最高のシングルシードベースラインに対して64%のばらつき違反を減じるが、あらゆる検出器の4シードアンサンブルはマージンを0.01に絞る(一致したカバレッジで区別できない$p=0.21$)。
クロスモデル転送は、内部からクロスファミリーへ、オープンウェイトからAPIへ、部分的にはマルチモデルトレーニングによって、低下する。
オーディエンスのためのフレームワークを提供する。本当の障壁はディテクターの分散であり、分散シフトではない。
関連論文リスト
- Auditing Routing Entropy as an Uncertainty Signal in Attention-Residual Transformers [14.727396825167917]
We audit that reading for routing entropy in Attention-Residual (AR)variants of Swin-Tiny and DeiT-Small。
我々は、このトレースが、モデル自身の自信が明らかにしている以上の正確性に関する情報を持っているかどうかを問う。
論文 参考訳(メタデータ) (2026-10-01T11:37:19Z) - DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text [42.255101883505056]
DeBERTa-ConParaは,HC3 Plus,M4,MAGE,RAIDで学習した文脈変換器エンコーダと,攻撃対応Unicode前処理を組み合わせたデプロイメント指向検出器である。
我々の中心的な発見は、前処理が適用された場所によって反対方向に作用することである。
2つの配置の因子的変化は、正規化推論による生のトレーニングを最高の設定として独立に識別する。
12種類の攻撃クラスのうち、ホモグリーフとゼロ幅空間の挿入は11.05%から1.12%から96.98%に増加した。
論文 参考訳(メタデータ) (2026-10-01T00:57:33Z) - Differentiable Interval Bottlenecks for Interpretable Anomaly Detection in Numerical Data [0.5213778368155993]
DIFFINTは、遅延ボトルネックがソフトな軸方向のインターバルメンバシップのセットとして構成されるオートエンコーダである。
統計的にタイトされた7つの方法のリードクラスターの中で唯一の解釈可能な検出器である。
論文 参考訳(メタデータ) (2026-09-03T14:05:27Z) - When Verifiers Vote Backwards under Verdict Substitution: Signed Pivotal Value in Correlated Self-Consistency [9.710464466895521]
1回の投票で決定されるクエリのみ 多数決を変更できる。
ラベル付き検証スタイルの介入を用いて,その変化の兆候について検討する。
論文 参考訳(メタデータ) (2026-08-08T05:02:34Z) - Post-Hoc Trajectory-Risk Certification for Modular LLM-Based Security Agents [1.4915002678801434]
ボンフェロニ配置は分布自由であるが,相関誤差下では保守的であることを示す。
粗いラベルの選択は、学習された依存なしにほぼ完璧な相関関係を作ることができる。
ステージ単位の証明書とペアの重なり合うバウンドを用いたモジュール証明書は、正の平均利得が0.6%となる。
論文 参考訳(メタデータ) (2026-08-04T23:48:36Z) - Level, Sharpness, and Corpus: Why Zero-Shot OOD Detector Rankings Do Not Transfer [36.02174685737611]
各検出器は、少なくとも1つのドメインで80%$ FPR95 を超え、所望の検出器は、分配データと基礎となる VLM の両方に依存する。
本稿では, ベース検出器, レベル, シャープネスの非補償融合により, 補完的証拠を保存する検出器非依存ラッパーであるComplementary Evidence Guard(CEG)を紹介する。
論文 参考訳(メタデータ) (2026-07-29T08:00:49Z) - Consistent Evidence, Robust Recognition: Faithful Attribution Regularization under Geometric Transformations [42.90321348046055]
画像領域のサブモジュール探索に基づくアノテーションのない属性正規化フレームワークを提案する。
候補部分集合がモデル出力にどのように影響するかを測定することで、探索は、コンパクトでクラス識別的な証拠を探索由来の監督として抽出する。
ImageNet-100の実験結果から,VT-B/16の属性安定性,挿入性,削除性は0.28ポイントの精度低下で大幅に向上することがわかった。
論文 参考訳(メタデータ) (2026-07-26T20:41:38Z) - Unveiling the Entropy Dynamics of Chain-of-Thought Reasoning [81.57028614960576]
1)信頼性の高い -- 信頼性の高い -- 信頼性の高い領域での回答は高度に正確で安定したものになり、2)高冗長性 -- モデルは正しい回答に達した後ずっと経ってから不必要なトークンを生成する。
これらの特性はより効率的で信頼性の高い推論戦略を解き放つ。
論文 参考訳(メタデータ) (2026-06-01T10:11:14Z) - Distributional Energy-Based Models for Uncertainty-Aware Structured LLM Reasoning [40.342912574072024]
大規模言語モデルは、旅行計画やコードソリューションのような構造化されたアウトプットを生成する。
個々の推論ステップは正しく見えるが、アウトプット全体が予算に違反したり、テストケースに失敗したり、あるいは以前の推論に矛盾することがある。
構造化LCM出力の検証のための決定論的解析制約付き学習品質スコアラを提案する。
論文 参考訳(メタデータ) (2026-05-15T17:08:27Z) - Probabilistic Soundness Guarantees in LLM Reasoning Chains [37.440902632372904]
ARES(Autoregressive Reasoning Entailment Stability)は、事前に検証された前提のみに基づいて、各推論ステップを評価する確率的フレームワークである。
ARESは4つのベンチマークで最先端のパフォーマンスを達成し、非常に長い合成推論チェーン上で優れた堅牢性を示す。
論文 参考訳(メタデータ) (2025-07-17T09:40:56Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。