論文の概要: Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring
- arxiv url: http://arxiv.org/abs/2604.18835v1
- Date: Mon, 20 Apr 2026 20:59:25 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-04-22 22:41:49.490959
- Title: Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring
- Title(参考訳): 文書ヘイスタックにおける意味的ニーズ:LCM-as-a-Judge類似性検査の感度試験
- Authors: Sinan G. Aksoy, Alexandra A. Sabrio, Erik VonKaenel, Lee Burke,
- Abstract要約: 文書比較の微妙な意味変化に対してLLM感度を探索するスケーラブルな多要素実験フレームワークを提案する。
LLMは文書内位置バイアスを呈するなど,いくつかの顕著な知見が得られた。
提案フレームワークは,現在および将来のモデル間でのスコアリング動作の監査と比較を行うための,実用的でLCMに依存しないツールキットを提供する。
- 参考スコア(独自算出の注目度): 39.19569497931068
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: We propose a scalable, multifactorial experimental framework that systematically probes LLM sensitivity to subtle semantic changes in pairwise document comparison. We analogize this as a needle-in-a-haystack problem: a single semantically altered sentence (the needle) is embedded within surrounding context (the hay), and we vary the perturbation type (negation, conjunction swap, named entity replacement), context type (original vs. topically unrelated), needle position, and document length across all combinations, testing five LLMs on tens of thousands of document pairs. Our analysis reveals several striking findings. First, LLMs exhibit a within-document positional bias distinct from previously studied candidate-order effects: most models penalize semantic differences more harshly when they occur earlier in a document. Second, when the altered sentence is surrounded by topically unrelated context, it systematically lowers similarity scores and induces bipolarized scores that indicate either very low or very high similarity. This is consistent with an interpretive frame account in which topically-related context may allow models to contextualize and downweight the alterations. Third, each LLM produces a qualitatively distinct scoring distribution, a stable "fingerprint" that is invariant to perturbation type, yet all models share a universal hierarchy in how leniently they treat different perturbation types. Together, these results demonstrate that LLM semantic similarity scores are sensitive to document structure, context coherence, and model identity in ways that go beyond the semantic change itself, and that the proposed framework offers a practical, LLM-agnostic toolkit for auditing and comparing scoring behavior across current and future models.
- Abstract(参考訳): 文書比較において,LLMの感度を微妙な意味的変化に対して体系的に探索する,スケーラブルで多要素的な実験フレームワークを提案する。
一つの意味論的に変化した文(針)が周囲の文脈(干し草)に埋め込まれており、摂動タイプ(ネゲーション、接続スワップ、名前付きエンティティ置換)、コンテキストタイプ(元来は無関係)、針の位置、文書の長さが異なる。
分析の結果,いくつかの顕著な所見が得られた。
まず、LCMは、以前に研究された候補順序効果とは異なる文書内位置バイアスを示す。
第二に、変化した文が位相的に無関係な文脈に囲まれている場合、体系的に類似度を下げ、非常に低いか非常に高い類似度を示す双極化スコアを誘導する。
これは、トポロジ的に関連するコンテキストによって、モデルがコンテキスト化および重み付けを行うことができる解釈的フレームアカウントと一致している。
第三に、各 LLM は定性的に異なるスコアリング分布を生成し、安定な「フィンガープリント」は摂動型に不変であるが、全てのモデルは摂動型をいかに寛大に扱うかについて普遍的な階層を共有している。
これらの結果から,LLMの意味的類似度スコアは,意味的変化そのものを超越した方法で,文書構造やコンテキストコヒーレンス,モデル識別に敏感であり,提案フレームワークは,現在および将来のモデルにおける評価行動の監査・比較を行うための,実用的かつLLMに依存しないツールキットを提供することを示した。
関連論文リスト
- HCRE: LLM-based Hierarchical Classification for Cross-Document Relation Extraction with a Prediction-then-Verification Strategy [54.91468501159335]
文書間関係抽出 (RE) は, 異なる文書に存在する頭部尾部エンティティ間の関係を識別することを目的としている。
本稿では,各レベルでの多視点検証により信頼性を向上させる推論戦略を提案する。
論文 参考訳(メタデータ) (2026-04-09T07:55:27Z) - UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning [101.62386137855704]
本稿では,Universal Multimodal Embedding (UniME-V2)モデルを提案する。
提案手法はまず,グローバル検索による潜在的な負のセットを構築する。
次に、MLLMを用いてクエリ候補対のセマンティックアライメントを評価するMLLM-as-a-Judge機構を提案する。
これらのスコアは、ハード・ネガティブ・マイニングの基礎となり、偽陰性の影響を緩和し、多様な高品質なハード・ネガティブの識別を可能にする。
論文 参考訳(メタデータ) (2025-10-15T13:07:00Z) - When LLMs Disagree: Diagnosing Relevance Filtering Bias and Retrieval Divergence in SDG Search [0.0]
大規模言語モデル(LLM)は、情報検索パイプラインに文書関連ラベルを割り当てるのにますます使われている。
LLMは境界線のケースにしばしば反対し、そのような不一致が下流の検索にどのように影響するかという懸念を提起する。
モデル不一致は体系的であり、ランダムではないことを示す。
本稿では,検索評価における分析対象として分類不一致を用いることを提案する。
論文 参考訳(メタデータ) (2025-07-02T20:53:51Z) - QUDsim: Quantifying Discourse Similarities in LLM-Generated Text [70.22275200293964]
本稿では,会話の進行過程の違いの定量化を支援するために,言語理論に基づくQUDと質問意味論を紹介する。
このフレームワークを使って$textbfQUDsim$を作ります。
QUDsimを用いて、コンテンツが異なる場合であっても、LLMはサンプル間で(人間よりも)談話構造を再利用することが多い。
論文 参考訳(メタデータ) (2025-04-12T23:46:09Z) - Revisiting Word Embeddings in the LLM Era [0.2999888908665658]
大規模言語モデル(LLM)は、最近、様々なNLPタスクにおいて顕著な進歩を見せている。
従来の非コンテクスト化単語と文脈化単語の埋め込みをLLMによる埋め込みで比較した。
以上の結果から,LLMは意味的関連語をより緊密にクラスタ化し,非文脈化設定における類似処理をより良く行うことが示唆された。
論文 参考訳(メタデータ) (2025-02-26T22:45:08Z) - Estimating Commonsense Plausibility through Semantic Shifts [66.06254418551737]
セマンティックシフトを測定することでコモンセンスの妥当性を定量化する新しい識別フレームワークであるComPaSSを提案する。
2種類の細粒度コモンセンス可視性評価タスクの評価は,ComPaSSが一貫してベースラインを上回っていることを示している。
論文 参考訳(メタデータ) (2025-02-19T06:31:06Z) - Revisiting Word Embeddings in the LLM Era [5.122866382023337]
大規模言語モデル(LLM)は、最近、様々なNLPタスクにおいて顕著な進歩を見せている。
従来の非コンテクスト化単語と文脈化単語の埋め込みをLLMによる埋め込みで比較した。
以上の結果から,LLMは意味的関連語をより緊密にクラスタ化し,非文脈化設定における類似処理をより良く行うことが示唆された。
論文 参考訳(メタデータ) (2024-02-16T21:47:30Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。