論文の概要: Sixteen models, fewer than two voices: measuring ensemble dispersion where no answer is uniquely correct
- arxiv url: http://arxiv.org/abs/2608.00285v1
- Date: Fri, 31 Jul 2026 20:42:53 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-04 15:07:24.566425
- Title: Sixteen models, fewer than two voices: measuring ensemble dispersion where no answer is uniquely correct
- Title(参考訳): 16モデル、2声未満:答えが一意に正しくないアンサンブル分散の測定
- Authors: Mario Vega-Barbas, Lidia Mora-Valenciano, Iván Pau, Fernando Seoane, Farhad Abtahi,
- Abstract要約: 平均して10の家系から抽出された16の言語モデルでは、精神療法ケースの1.69の異なる定式化のセマンティック多様性が生み出された。
アウトプットの分散は多様性と不確実性の両方として測定される。
- 参考スコア(独自算出の注目度): 37.64799781700206
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Sixteen language models drawn from ten families produced, on average, the semantic diversity of 1.69 distinct formulations of a psychotherapeutic case, against a single-model baseline of 1.43 from one model's own runs. Ensembles place more than one reading before a decision-maker on the premise that several models supply several perspectives. Dispersion over their outputs is measured both as diversity and as uncertainty, and both traditions validate it against a correctness criterion that this task does not admit. Measuring diversity is a solved problem: the Vendi Score, the exponential of the von Neumann entropy of a similarity matrix, is an effective number of distinct elements. What a single aggregate does not say is where the diversity comes from. We define a per-model dissent contribution, the complement of a model's mean similarity to the other members of its ensemble: a magnitude from the same matrix, not a decomposition of the spectral index, whose maximum identifies the most divergent voice. Crossing model and case, we test as a preregistered hypothesis whether model identity accounts for a non-zero share of the variance in dissent, and characterise the structure that test detects. The panel formulated fifteen stratified vignettes, yielding 7,082 formulations for analysis. Model identity was a detectable structuring factor of the dissent that remained, but the usual categories recovered it only partly: scale differences pointed in opposite directions across pairs, family grouped models on only five two-member lines, and the most divergent voice changed with panel composition, so that the surfaced outlier describes the ensemble rather than the model. Dissent did not track the interpretive openness for which the case bank was stratified; it was organised by clinical content instead, leaving the dispersion an ensemble produces a property to measure rather than assume.
- Abstract(参考訳): 16の言語モデルは、平均して、精神療法ケースの1.69のセマンティックな多様性を、あるモデル自身の実行から1.43の単一モデルベースラインに対して生成した。
アンサンブルは、複数のモデルがいくつかの視点を提供するという前提で、意思決定者の前に複数の読み物を置く。
アウトプットの分散は多様性と不確実性の両方として測定され、両方の伝統は、このタスクが認めない正確性基準に対してそれを検証している。
類似性行列のフォン・ノイマンエントロピーの指数であるヴェンディスコア(Vendi Score)は、異なる要素の有効数である。
一つの集合が言っていないのは、多様性がどこから来ているかだ。
モデル毎の不一致の寄与(モデルの平均的類似度)を定義し、そのアンサンブルの他のメンバーとの平均的類似度(スペクトル指数の分解ではなく、同じ行列からの等級)を定義し、最も発散した音声を最大で識別する。
モデルとケースを交差させて、モデルアイデンティティが不一致の分散の非ゼロシェアを占めるかどうかを予め登録された仮説として検証し、テストが検出する構造を特徴づける。
パネルは15個の層状ビグネットを定式化し、解析のために7,082個の定式化を行った。
モデルの同一性は、不一致の検知可能な構造因子であったが、通常のカテゴリでは、ペア間で反対方向を指し示すスケール差、わずか5つの2つのメンバーライン上のファミリーグループモデル、パネル構成によって最も異なる音声が変化したため、表面のアウトリアはモデルではなくアンサンブルを記述する。
反対派は、ケースバンクが階層化された解釈的開放性を追跡せず、代わりに臨床内容によって組織化され、分散をアンサンブルに残して、仮定するよりも測定する特性を生み出した。
関連論文リスト
- CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models [3.991518853474585]
我々は,認知タスクが保証次元ラベルをいつ得点するかを決定するためのベンチマークであるCogArenaを紹介する。
55のオープンウェイトモデル全体において、ほぼすべてのパラダイム相関は正であり、共通の軸が約半分の分散を説明する。
グループ内のアドバンテージは小さく、スコアリングに敏感で、モデルファミリ全体にわたって不確実である。
論文 参考訳(メタデータ) (2026-07-27T18:56:13Z) - Traits Run Deeper: Trait-Specific Asymmetric Fusion for Personality Assessment [53.50824432699408]
Traits Run Deeperは、新しいパーソナリティアセスメントフレームワークである。
MFR(Multimodal Foundation Representation)、TSMF(Trit-Specific Modality Fusion)、DCPR(Distributed-Calibrated Personality Regression)の3つのコンポーネントで構成されている。
論文 参考訳(メタデータ) (2026-06-09T06:38:36Z) - LLMs Show No Signs Of Individuated Metacognition [0.023227405857540805]
20大言語モデルから二項信頼判断を分解する。
信頼性が異なる2つのモデルも性能が異なるかどうかを問う。
いずれの検査領域においても,有意な弁別メタ認知の証拠は見つからない。
論文 参考訳(メタデータ) (2026-05-22T23:54:33Z) - A Mechanistic Study of Tabular Foundation Models [12.26738662548555]
タブラル基礎モデルは、様々な分類タスクと回帰タスクに精度で収束する。
これは、リーダーボードが答えられない疑問を引き起こす。
i)モデルが同一のコンテキスト内アルゴリズムを実行するかどうか,(ii)行,列,およびクラス置換不変性が生じるか,および(iii)モデルが推論されたメカニズムに対してエンジニアリングされた摂動下にあるか,の3つを特徴付ける。
論文 参考訳(メタデータ) (2026-05-20T15:23:16Z) - Identification of Causal Direction under an Arbitrary Number of Latent Confounders [54.76982125821112]
実世界のシナリオでは、観測された変数は複数の潜伏変数によって同時に影響を受けることがある。
我々は,特定の方法で構築された観測変数の高次累積行列を併用する。
これらの高次累積行列の階数不足特性から,2つの観測変数間の因果非対称性が直接観察可能であることを示す。
論文 参考訳(メタデータ) (2025-10-26T15:10:00Z) - Counting Like Human: Anthropoid Crowd Counting on Modeling the
Similarity of Objects [92.80955339180119]
メインストリームの群衆計数法は 密度マップを補強して 計数結果を得るために統合する。
これに触発された我々は,合理的かつ人為的な集団カウントフレームワークを提案する。
論文 参考訳(メタデータ) (2022-12-02T07:00:53Z) - Causal Discovery in Linear Latent Variable Models Subject to Measurement
Error [29.78435955758185]
線形系における測定誤差の存在下での因果発見に着目した。
我々は、この問題と因果発見の驚くべき関連性を、観察されていない親性原因の存在で示している。
論文 参考訳(メタデータ) (2022-11-08T03:43:14Z) - Predicting Out-of-Domain Generalization with Neighborhood Invariance [59.05399533508682]
局所変換近傍における分類器の出力不変性の尺度を提案する。
私たちの測度は計算が簡単で、テストポイントの真のラベルに依存しません。
画像分類,感情分析,自然言語推論のベンチマーク実験において,我々の測定値と実際のOOD一般化との間に強い相関関係を示す。
論文 参考訳(メタデータ) (2022-07-05T14:55:16Z) - Learning Disentangled Representations with Latent Variation
Predictability [102.4163768995288]
本稿では,潜在不整合表現の変動予測可能性について述べる。
逆生成プロセス内では、潜時変動と対応する画像対の相互情報を最大化することにより、変動予測可能性を高める。
本研究では,潜在表現の絡み合いを測るために,基礎的構造的生成因子に依存しない評価指標を開発する。
論文 参考訳(メタデータ) (2020-07-25T08:54:26Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。