論文の概要: The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure
- arxiv url: http://arxiv.org/abs/2608.25005v1
- Date: Tue, 25 Aug 2026 18:00:18 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-27 14:15:15.380774
- Title: The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure
- Title(参考訳): 大規模言語モデルでは不完全なパラドックスは必要ではない:モデルが失敗する前にベンチマーク失敗
- Abstract要約: 本稿では,ベンチマークと結論について検討し,概念的および評価的ミス・スペクテーションに大きく影響していることを示す。
厳格なNLI標準の下では、グループAのインスタンスの76%は、決定を明示的に除外していない。
イベントセマンティックなNLIを多段階推論問題として定式化し、中間的意味決定と最終的な予測の両方を評価する。
- 参考スコア(独自算出の注目度): 42.25959731143407
- License: http://creativecommons.org/licenses/by-nc-nd/4.0/
- Abstract: The imperfective paradox provides a useful test of compositional semantic analysis. Recent work constructs an NLI benchmark and reports that models frequently infer completed telic events from progressive descriptions, attributing this behavior to a Teleological Bias. It further argues that prompting interventions cause a Calibration Crisis. We reexamine the benchmark and conclusions and show that it is substantially affected by conceptual and evaluation mis-specifications. We identify three conceptual mis-specifications. In particular, Aspectual Reduction affects the benchmark construction, analysis, experiments, and conclusions. Under a strict NLI standard, 76% of Group A instances do not explicitly rule out culmination. In our native-speaker annotation, 38% of Group A examples and 29% of the Group C examples were judged to permit an alternative interpretation. To control these issues and lexical variation, we construct Lexically Matched Minimal Pairs. At the evaluation level, we formulate event-semantic NLI as a Multi-step Reasoning Problem and assess both intermediate semantic decisions and final predictions. Our results show that models often do not affirm culmination but nevertheless accept the corresponding simple-past hypothesis, a pattern we characterize as Sufficiency Bias. We further show that prompting interventions produce a Decision Shift among labels without reliably improving the underlying semantic understanding and reasoning. Intermediate and oracle-guided analyses identify two additional failure modes: errors in compositional aspectual classification and Surface-form Attraction toward surface-associated answers. Our experiments on Qwen-7B with suitable prompts, GPT-5.4, and Qwen-72B provide initial evidence for the context sensitivity of aspectual classification and suggest that these models can achieve performance comparable to that of human annotators.
- Abstract(参考訳): 不完全なパラドックスは、構成意味解析の有用なテストを提供する。
最近の研究は、NLIベンチマークを構築し、モデルがプログレッシブな記述から完了したテリック事象をしばしば推測し、この振る舞いがテレロジカルバイアスに起因することを報告している。
さらに、介入を促すことが校正危機を引き起こすと主張している。
我々は、ベンチマークと結論を再検討し、概念的および評価的誤特定の影響を著しく受けていることを示す。
3つの概念的誤解を識別する。
特に、アスペクトリダクションは、ベンチマークの構築、分析、実験、結論に影響を与える。
厳格なNLI標準の下では、グループAのインスタンスの76%は、決定を明示的に除外していない。
ネイティブスピーカーアノテーションでは,グループA例の38%,グループC例の29%が代替解釈を許容すると判断された。
これらの問題と語彙変化を制御するため、Lexically Matched Minimal Pairsを構築した。
評価レベルでは、イベントセマンティックなNLIを多段階推論問題として定式化し、中間的意味決定と最終的な予測の両方を評価する。
以上の結果から,モデルが決定を肯定しないことが多いが,それにもかかわらず,我々は十分バイアスと特徴づけるパターンである単純最小仮説を受け入れていることがわかった。
さらに,介入を促すことでラベル間の決定的シフトが生じ,その基盤となる意味理解や推論を確実に改善することができないことを示す。
中間およびオラクル誘導分析では、組成的側面分類における誤差と表面形状の抽出という2つの障害モードが同定される。
適切なプロンプトを持つQwen-7B, GPT-5.4, Qwen-72Bを用いた実験により, アスペクト的分類の文脈感度に関する最初の証拠が得られ, これらのモデルが人間のアノテータに匹敵する性能を達成できることが示唆された。
関連論文リスト
- Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery [95.44432473367836]
大規模言語モデル(LLM)は、事前に特定された質問に答える能力に優れるが、未解決の未解決段階をナビゲートする能力はほとんど測定されていない。
仮説探索(PHD)を導入し,不確定な証拠から,仮説空間を自律的に構築することをモデルに求める。
本稿では,6つの科学的・分析的領域にわたる988ケースのベンチマークであるHypoArenaと,オープンエンド仮説セットの評価フレームワークであるHypoEvalを紹介する。
論文 参考訳(メタデータ) (2026-07-17T08:56:43Z) - Aligning LLM Uncertainty with Human Disagreement in Subjectivity Analysis [57.78407973423517]
我々は、モデルが人間の不一致を反映した不確実性を表現しながら予測する不確実性を考慮した主観性分析を提唱する。
この視点を運用するために,二相認識と不確実性アライメントの枠組みを提案する。
3つの主観的分析課題の実験は、DPUAが人間の不一致とモデルの不確実性をよりよく整合させながら、タスク性能を保っていることを示している。
論文 参考訳(メタデータ) (2026-05-11T11:52:58Z) - Deconfounding Scores and Representation Learning for Causal Effect Estimation with Weak Overlap [75.41420283374737]
オーバーラップは因果治療効果推定の鍵となる条件である。
本稿では,デコンウンディングスコア(deconfounding scores)と呼ばれる特徴表現のクラスを提案する。
このクラスでは, 予後スコアが重なり合うことを示す。
論文 参考訳(メタデータ) (2026-04-01T12:19:42Z) - Measuring Language Model Hallucinations Through Distributional Correctness [7.106986689736826]
この問題を解決するために,新しい評価基準である分布補正スコア(DCS)を導入した。
DCSは、誤った回答における有害な過信と、棄権によって表される不確実性を区別し、解釈可能なデフォルト範囲でスコアを提供する。
DCSは、推測よりも真に不確実性を表現するモデルにインセンティブを与える、よりニュアンスで整列した評価パラダイムを提供する。
論文 参考訳(メタデータ) (2025-10-05T17:50:42Z) - A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models [58.32070787537946]
思考の連鎖(CoT)推論は、大きな言語モデルの性能を高める。
大規模視覚言語モデルにおけるCoT忠実度に関する最初の総合的研究について述べる。
論文 参考訳(メタデータ) (2025-05-29T18:55:05Z) - Estimating the Causal Effects of Natural Logic Features in Transformer-Based NLI Models [16.328341121232484]
文脈介入の効果を測定するために因果効果推定手法を適用した。
本研究はトランスフォーマーの無関係な変化に対する堅牢性と影響の高い変化に対する感受性について検討する。
論文 参考訳(メタデータ) (2024-04-03T10:22:35Z) - HANS, are you clever? Clever Hans Effect Analysis of Neural Systems [1.6267479602370545]
大規模言語モデル(It-LLM)は、認知状態、意図、そしてすべての人々の反応を推論する優れた能力を示しており、人間は日々の社会的相互作用を効果的にガイドし理解することができる。
モデル能力の確固たる評価を構築するために、MCQ(Multiple-choice Question)ベンチマークがいくつか提案されている。
しかし、初期の研究は、I-LLMに固有の「順序バイアス」があることを示しており、適切な評価に挑戦している。
論文 参考訳(メタデータ) (2023-09-21T20:52:18Z) - Achieving Equalized Odds by Resampling Sensitive Attributes [13.114114427206678]
等価性の概念をほぼ満足する予測モデルを学習するためのフレキシブルなフレームワークを提案する。
この微分可能な関数は、モデルパラメータを等化奇数に向けて駆動するペナルティとして使用される。
本研究は,予測規則が本性質に反するか否かを検出するための公式な仮説テストを開発する。
論文 参考訳(メタデータ) (2020-06-08T00:18:34Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。