論文の概要: Auditing Pairwise Equivalence Judgments: Self-Critique Effects and Diversity Measurement in Multi-Agent Hypothesis Generation
- arxiv url: http://arxiv.org/abs/2610.04133v1
- Date: Fri, 02 Oct 2026 23:07:36 GMT
- ステータス: 情報取得中
- システム内更新日: 2026-10-06 21:04:27.270153
- Title: Auditing Pairwise Equivalence Judgments: Self-Critique Effects and Diversity Measurement in Multi-Agent Hypothesis Generation
- Title(参考訳): Pairwise Equivalence Judgments:マルチエージェント仮説生成における自己批判効果と多様性の測定
- Abstract要約: 用語周波数-逆文書周波数(TF-IDF)の類似性、密埋め込み、LLM-as-a-judgeの3つの等価ルールを比較した。
3つのルールはいずれも意味保存編集に不変であるが、開始イベントが置き換えられると、LSM-as-a-judgeは有効なペアの83%を異なるメカニズムとして識別する。
- 参考スコア(独自算出の注目度): 0.8517712901743845
- License:
- Abstract: Multi-agent systems built on large language models (LLMs) are increasingly applied to scientific discovery and hypothesis generation. Both the effect of refinement and the diversity of the delivered set are hard to interpret before experimental ground truth exists, and both are typically reported by deciding whether pairs of generated hypotheses describe the same underlying mechanism. We study two evaluation questions that rest on this pairwise equivalence judgment: (1) how much self-critique changes delivered hypotheses beyond run-to-run variability, and (2) how the equivalence rule used to group hypotheses affects measured diversity. Across four proprietary instances, we hold opening hypotheses fixed, rerun the downstream workflow with 0, 1, and 5 critique rounds, and score matched hypothesis pairs with an LLM-as-a-judge. Relative to matched same-depth reruns, moving from 0 to 1 round produces 34.5 percentage points (pp) of additional mechanism-level divergence, whereas 1 to 5 rounds adds 1.3 pp. We then compare three equivalence rules: term frequency--inverse document frequency (TF--IDF) similarity, dense embeddings, and the same LLM-as-a-judge. We construct controlled hypothesis pairs that either preserve the causal explanation through wording or biological-terminology changes, or replace one component of the causal chain while holding the rest fixed. All three rules are invariant to meaning-preserving edits, but when the initiating event is replaced, the LLM-as-a-judge identifies 83% of valid pairs as different mechanisms, versus 0% for TF--IDF and 8% for embeddings; varying only the rubric that defines same mechanism moves this figure from 38% to 96%. Together, these results show that pairwise equivalence judgments are a measurement choice: how mechanism equivalence is defined affects both the estimated effect of self-critique and the measured diversity of generated hypotheses.
- Abstract(参考訳): 大規模言語モデル(LLM)上に構築されたマルチエージェントシステムは、科学的発見と仮説生成にますます応用されている。
改良の効果と納品された集合の多様性は、実験的な基底真理が存在する前に解釈することは困難であり、両者は典型的には、生成された仮説のペアが同じ基盤メカニズムを記述しているかどうかを判断することによって報告される。
1) 自己批判的変化が, 実行時から実行時までの変動性を超えた仮説をどの程度提示するか, (2) 仮説をまとめるために用いられる同値性ルールが, 測定された多様性にどのように影響するか, という2つの評価問題について検討した。
4つのプロプライエタリなインスタンスで、オープニング仮説を固定し、0, 1, 5の批判ラウンドで下流のワークフローを再実行し、LLM-as-a-judgeと一致した仮説ペアをスコア付けします。
一致した同一深度の再走に対して、0から1ラウンドに移動すると、追加の機構レベルの分岐点が34.5ポイント (pp) 、一方、1から5ラウンドは1.3 pp となる。
次に、項周波数-逆文書周波数(TF-IDF)類似性、密埋め込み、LLM-as-a-judgeの3つの等価規則を比較した。
本研究は,口頭弁別や生物学的終末変化による因果的説明を維持するか,あるいは残余を固定したまま因果的連鎖の1成分を置き換える制御された仮説ペアを構築した。
3つのルールはいずれも意味保存編集に不変であるが、開始イベントが置き換えられた場合、LSM-as-a-judgeは有効なペアの83%を異なるメカニズムとみなし、TF-IDFは0%、埋め込みは8%と定義している。
これらの結果は、ペアワイズ等価性判定が測定選択であることを示し、メカニズム等価性の定義が自己批判の見積効果と生成された仮説の多様性にどのように影響するかを示す。
関連論文リスト
- The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure [42.25959731143407]
本稿では,ベンチマークと結論について検討し,概念的および評価的ミス・スペクテーションに大きく影響していることを示す。
厳格なNLI標準の下では、グループAのインスタンスの76%は、決定を明示的に除外していない。
イベントセマンティックなNLIを多段階推論問題として定式化し、中間的意味決定と最終的な予測の両方を評価する。
論文 参考訳(メタデータ) (2026-08-25T18:00:18Z) - Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate [5.650336594658653]
従来の解答フリップは, 自発的不安定性, 姿勢による適合性, 推論による説得の3つのメカニズムを混同している。
我々の3ソース分解フレームワークは、制御された対策条件によってそれぞれを分離する。
論文 参考訳(メタデータ) (2026-05-30T17:41:11Z) - How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles [46.63622714488747]
共有事前学習データ、蒸留、アライメントパイプラインは、隠れた振る舞い依存、潜伏絡みを誘導することができる。
実際には、これは相関した推論パターンと同期された障害として現れます。
ブラックボックス言語モデル間の行動絡みを監査するための統計的枠組みを開発する。
論文 参考訳(メタデータ) (2026-04-08T23:32:06Z) - When Are Learning Biases Equivalent? A Unifying Framework for Fairness, Robustness, and Distribution Shift [0.0]
機械学習システムは、保護されたグループに対する不公平さ、刺激的な相関に対する脆さ、少数民族のサブ人口に対するパフォーマンスの低下など、さまざまな障害モードを示す。
本稿では,異なるバイアス機構がモデル性能に定量的に等価な効果をもたらすことを特徴付ける統一理論フレームワークを提案する。
論文 参考訳(メタデータ) (2025-11-09T20:48:09Z) - Statistical Hypothesis Testing for Auditing Robustness in Language Models [49.1574468325115]
本稿では,摂動解析を頻繁な仮説テスト問題として再検討するフレームワークである分布に基づく摂動解析を紹介する。
モンテカルロサンプリングを用いて低次元意味的類似性空間内に経験的ヌルおよび代替出力分布を構築する。
反応変化の定量化、正/偽の正率の測定、参照モデルとの整合性の評価について述べる。
論文 参考訳(メタデータ) (2025-06-09T17:11:07Z) - Noninterference Analysis of Irreversible or Reversible Systems with Nondeterminism and Probabilities [42.52342528033571]
非干渉理論は、マルチレベルセキュリティシステムにおけるセキュアな計算の分析をサポートする。
非決定論的設定では、弱い双相似性を通して非干渉を評価することは可逆的システムには適しているが、可逆的双相似性に分岐するシステムに対しては、最近より適切であることが証明されている。
我々は、可逆系と可逆系のそれぞれに弱および分岐双相似の確率的不変量を採用することにより、非干渉特性を再送する。
論文 参考訳(メタデータ) (2025-01-31T16:49:42Z) - Simultaneous inference for generalized linear models with unmeasured confounders [0.0]
本稿では,構造を利用して線形射影を3つの重要な段階に統合する,統一的な統計的推定と推測の枠組みを提案する。
サンプルおよび応答サイズとして$z$-testsの効果的なType-Iエラー制御が無限大に近づくことを示す。
論文 参考訳(メタデータ) (2023-09-13T18:53:11Z) - Nested Counterfactual Identification from Arbitrary Surrogate
Experiments [95.48089725859298]
観測と実験の任意の組み合わせからネスト反事実の同定について検討した。
具体的には、任意のネストされた反事実を非ネストされたものへ写像できる反ファクト的非ネスト定理(英語版)(CUT)を証明する。
論文 参考訳(メタデータ) (2021-07-07T12:51:04Z) - Correct block-design experiments mitigate temporal correlation bias in
EEG classification [68.85562949901077]
[1]の主主張は極めて過大評価されており、他の分析は間違った方法論的選択によって深刻な欠陥を負っていることを示す。
脳波の時間相関が2つの実験環境で同じモデルをテストすることによって分類精度に及ぼす影響について検討した。
論文 参考訳(メタデータ) (2020-11-25T22:25:21Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。