論文の概要: Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit
- arxiv url: http://arxiv.org/abs/2608.27309v2
- Date: Wed, 02 Sep 2026 05:16:14 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-03 15:35:44.738195
- Title: Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit
- Title(参考訳): 検閲型評価尺度における差分差は効果を生ずる:事前登録によるLCM-Judge監査からの証拠
- Abstract要約: LLM審査員の聴取は、一致した条件を対比することでバイアスを認定する。
このエンドポイントは、報告するスケールでは特定されない。
我々は,そのメカニズムをクローズドな形で導き,その貢献が監査の自己評価から測定可能であることを示す。
- 参考スコア(独自算出の注目度): 15.81584544074437
- License: http://creativecommons.org/licenses/by-nc-nd/4.0/
- Abstract: Audits of LLM judges certify a bias by contrasting matched conditions, and the strongest designs difference twice: a within-item contrast between two candidate responses, differenced again across a manipulated attribute, read off a bounded rating scale. We show that this endpoint is not identified on the scale that reports it. Each term of the double difference is censored by its own share, so the observed statistic confounds differential preference with differential attenuation: a severity shift common to both responses manufactures an interaction whenever the two censor it unequally, as unequal distances from the bounds make them, exactly where good stimuli place them. We exhibit the failure inside a pre-registered audit of a frozen pedagogy judge, sealed before the first of its 990 calls. The registered primary endpoint, the effect of a stated learner profile on the judge's scaffolding preference, is null: $+0.085$ points (95\% BCa $[-0.167, +0.353]$, $p = 0.684$). The audit's one nominally significant interaction, $+0.378$ ($p = 0.002$), is not identified as preference: a construction containing zero differential preference reproduces 79 to 85\% of it from the observed severity shift and the scale floor alone. We derive the mechanism in closed form and show that its contribution is measurable from an audit's own ratings.
- Abstract(参考訳): LLM審査員の聴取者は、一致した条件を対比することでバイアスを証明し、最も強い設計は2倍の差がある。
このエンドポイントは、報告するスケールでは特定されない。
両応答に共通する厳密なシフトは、境界からの不等距離がちょうど良い刺激が配置された場所にあるように、両者が不等間隔で検閲するたびに相互作用を発生させる。
990コールの初回前に封印された凍結教育審査員の事前登録監査の失敗を示す。
登録されたプライマリエンドポイントは、裁判官の足場設定に対する学習者プロファイルの影響である。$+0.085$ points (95\% BCa $[-0.167, +0.353]$, $p = 0.684$)。
オーディションの1つの名目上重要な相互作用である$+0.378$$(p = 0.002$)は、選好とは見なされない: ゼロ微分選好を含む構成は、観測された重度シフトとスケールフロアのみから、79から85%を再現する。
我々は,そのメカニズムをクローズドな形で導き,その貢献が監査の自己評価から測定可能であることを示す。
関連論文リスト
- Unmatched Does Not Mean False: Incomplete Reference Sets Can Reverse Calibration Rankings in Open-Ended Theory-of-Mind Tracking [0.5735035463793009]
Open-ended Theory-of-Mind (ToM) トラッカーは有限参照を欠く有効な信念を出力する。
有限参照+マーチャントパイプラインは、不正な出力をマークし、適切なスコアモデル選択をリバース可能なプロキシラベルを生成する。
論文 参考訳(メタデータ) (2026-08-26T11:36:31Z) - A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation [14.867869624585886]
評価器の構成妥当性を2次元プロファイルとして定式化する。
S と R は独立であり、スカラーの要約がすべての関連する比較を保存することはないことを示す。
これらの結果から, 評価対象構造の変化に対して, 高判定基準は弱い感度で共存できることが示唆された。
論文 参考訳(メタデータ) (2026-08-25T11:32:54Z) - Forged Peer Judgments Mislead Multimodal LLM Judge Panels: Source-Blind Anchoring and Panel-Consensus Verification [9.710464466895521]
独立した視覚的判断を求めることは、自己とピアフレーミングの両方の下で大きなアンカーリングギャップを生み出す。
本稿では,個別に収集した盲票に対する引用を相互にチェックするパネル合意検証を導入する。
論文 参考訳(メタデータ) (2026-08-08T04:55:45Z) - Visual Credit Audit for Multimodal Spatial Reasoning [70.16915309526443]
Visual Credit Auditは、ベンチマーク画像がテキストのみとブランクコントロールよりもモデルの宣言された決定をもっとサポートするかどうか、モデルが関係性固有の視覚的エビデンスに反応するかどうかの2つの評価を分離する。
ラベルを適用すれば依存性認定正当性(D-CC)が得られる
4つのオープンMLLMと2つの空間ベンチマーク、12.73-26.25%の判定は正確であるが、証明されていない。
論文 参考訳(メタデータ) (2026-07-29T15:55:31Z) - Auditing Belief-Conditioned LLM Agents in Hidden-Information Social Deduction Games [50.880420636090896]
9-player Werewolf環境において,隠れた役割に対する外部信頼状態を維持するための監査可能なフレームワークを構築した。
我々は,その効果を関連づけとして報告し,そのメカニズムを未解決として扱う。
論文 参考訳(メタデータ) (2026-07-12T16:03:30Z) - PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception [75.70273182182397]
PerceptionRubricsはルーブリックベースの評価フレームワークである。
飽和ベンチマークスコアと現実世界の脆さのギャップに対処する。
1,038枚のインフォメーションセンス画像と1万個以上のインスタンス固有のルーリックをペアリングする。
論文 参考訳(メタデータ) (2026-06-26T17:59:15Z) - A Pre-Registered Causal Partition of Self-Consistency Elicitation and Reward Design in RLVR [1.2958054117511815]
報酬からの強化学習は、報酬信号が刺激的であっても推論を改善する。
実践者は一般的に、報酬-設計効果として naive = acc(TRUE) - acc(R) を解釈する。
我々はこの推定が体系的に偏っていることを証明している。
論文 参考訳(メタデータ) (2026-06-04T09:35:54Z) - Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why [34.429649156970015]
最近のLLM-as-judge論文24件の調査では、判定尺度、ネクタイハンドリング、不正出力、禁断ハンドリングに絡み合ったメトリックの選択が見つかった。
Pearson's $r$、Spearman's $、Kendall's $_b$、phi係数$$、Matthews correlation Coefficientはすべて1つの数に還元される。
論文 参考訳(メタデータ) (2026-05-25T07:31:44Z) - A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness [57.510025257780306]
既存の検証プロトコルは、レッドチーム固有の分散シフトを考慮できないことを示す。
我々は、より一貫して判断可能な振る舞いのベンチマークであるReliableBenchと、判断失敗を公開するために設計されたデータセットであるJiceStressTestを提案する。
論文 参考訳(メタデータ) (2026-02-04T15:13:35Z) - TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them [58.04324690859212]
自動評価器(LLM-as-a-judge)としての大規模言語モデル(LLM)は、現在の評価フレームワークにおいて重大な矛盾を明らかにしている。
スコア比較不整合とペアワイズ・トランジティビティ不整合という2つの基本的不整合を同定する。
我々は2つの重要なイノベーションを通じてこれらの制限に対処する確率的フレームワークであるTrustJudgeを提案する。
論文 参考訳(メタデータ) (2025-09-25T13:04:29Z) - Mind the Gap: A Causal Perspective on Bias Amplification in Prediction & Decision-Making [58.06306331390586]
本稿では,閾値演算による予測値がS$変化の程度を測るマージン補数の概念を導入する。
適切な因果仮定の下では、予測スコア$S$に対する$X$の影響は、真の結果$Y$に対する$X$の影響に等しいことを示す。
論文 参考訳(メタデータ) (2024-05-24T11:22:19Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。