論文の概要: Lower-Resource, Higher Scores: Language Bias in LLM Evaluators
- arxiv url: http://arxiv.org/abs/2607.14480v2
- Date: Tue, 21 Jul 2026 06:28:06 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-22 14:48:36.199303
- Title: Lower-Resource, Higher Scores: Language Bias in LLM Evaluators
- Title(参考訳): 低リソース,高スコア:LLM評価器における言語バイアス
- Authors: Ej Zhou, Lucas Resck, Zheng Hui, Anna Korhonen,
- Abstract要約: 意味論的に同一の命令応答対を23言語で実験する。
その結果、多言語評価器は異なる評価言語に大きく異なるスコアを割り当てていることがわかった。
- 参考スコア(独自算出の注目度): 27.959960538188753
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: LLM evaluators (trained reward models and prompted LLM-as-a-Judge) are routinely validated via pairwise accuracy. In a multilingual setting, this operates under the premise that high pairwise accuracy implies reliable, language-neutral scoring. We show that this assumption does not hold. We conduct experiments with semantically identical instruction-response pairs across 23 languages, and find that multilingual evaluators assign significantly different scores to different evaluation languages. The bias is statistically significant and consistent across eight open-weight evaluators of different architectures and training paradigms, persists in frontier judges, and is strongly correlated with language resource level: lower-resource languages are scored more generously. Meanwhile, these biases are invisible to pairwise accuracy: evaluators achieve above 90% pairwise accuracy, yet have up to 43% difference in acceptance rate across languages under a global decision threshold, meaning, for instance, that harmful content in lower-resource languages is more likely to pass safety filters. Per-language thresholds would require language identification, which can be defeated by code-switched prompts. We then investigate why lower-resource languages receive higher rather than lower scores, and we find that model uncertainty is linked with the effect: models tend to give higher scores when less confident, both under negative log-likelihood and under token-free uncertainty measures; however, language identity remains a significant predictor after controlling for uncertainty, and the bias cannot be explained away by content difficulty alone, but is a structural, language-level misalignment.
- Abstract(参考訳): LLM評価器(トレーニングされた報酬モデルとLPM-as-a-Judge)は、ペアの精度で定期的に検証される。
多言語設定では、これは高いペアワイズ精度が信頼でき、言語中立のスコアリングを意味するという前提の下で機能する。
この仮定が成り立たないことを示す。
我々は23言語にまたがる意味論的に同一の命令応答対を用いて実験を行い、多言語評価器が異なる評価言語に異なるスコアを割り当てていることを見出した。
バイアスは統計的に有意であり、異なるアーキテクチャとトレーニングパラダイムの8つのオープンウェイト評価者の間で一貫性があり、フロンティアの判断に留まり、言語リソースレベルと強く相関している。
一方、これらのバイアスはペアの精度には見えない: 評価者はペアの精度を90%以上達成するが、グローバルな決定しきい値の下で言語間での受け入れ率に最大43%の差がある。
言語ごとのしきい値には言語識別が必要であり、コードに切り替えられたプロンプトによって破られる可能性がある。
次に,低リソース言語が低得点よりも高い結果を得る理由を考察し,モデルの不確実性が影響に結びついていることを見出した。負の対数的・無トークン不確実性の両面において,信頼度が低い場合には高いスコアを与える傾向にあるが,不確実性の制御後においても言語同一性は重要な予測因子であり,コンテンツ難易度だけではバイアスを説明できないが,構造的,言語レベルのミスアライメントである。
関連論文リスト
- Extending Item Response Theory for Efficient and Meaningful Multilingual Evaluation [19.965981631741894]
マルチ言語ベンチマークは、言語全体にわたる大規模言語モデル(LLM)の評価の中心である。
徹底的な評価は言語数とともに線形にスケールし、自動翻訳はスケールで見落とされるエラーを導入し、いくつかの項目は一般的な知識と文化固有の知識を詳述する。
我々はこれら3つを統一的な統計フレームワークであるMultilingual-IRTで解決し、言語ごとの難易度差による項目応答理論を拡張した。
論文 参考訳(メタデータ) (2026-06-14T07:16:03Z) - Shared Doubt: Zero-shot Cross-Lingual Confidence Estimation for Language Models [14.815594636224747]
中間表現から直接解の正しさを予測する軽量線形プローブを用いる。
学習されたレイヤの重みと複数の改善により、信頼性機能は言語全体の中間層に集中していることが明らかになった。
ゼロショットの言語間性能はソース言語と類似性に依存するが、プローブはリトレーニングなしで強力なベースラインを提供する。
論文 参考訳(メタデータ) (2026-05-29T12:25:24Z) - Investigating the Multilingual Calibration Effects of Language Model Instruction-Tuning [58.355275813623685]
本研究は,多言語設定における大規模言語モデル(LLM)の校正における重要なギャップについて考察する。
低リソース言語であっても、高リソース言語SFTデータセットのインストラクションチューニング後にモデルの信頼性が著しく向上する可能性がある。
しかし、精度の改善は限界的あるいは存在しないものであり、多言語言語における標準SFTの重大な欠点を浮き彫りにしている。
論文 参考訳(メタデータ) (2026-01-04T04:29:12Z) - Humans overrely on overconfident language models, across languages [32.71245803698373]
5言語にわたる多言語言語(ミス)校正,過信,過信のリスクについて検討した。
私たちの研究によると、言語全体で過度に信頼されるリスクが高いことが分かりました。
論文 参考訳(メタデータ) (2025-07-08T18:01:01Z) - PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts [85.78821098963607]
PolyMathは18の言語と4つの難易度をカバーする多言語数学的推論ベンチマークである。
我々のベンチマークは、包括性、言語多様性、高品質な翻訳の難しさを保証する。
論文 参考訳(メタデータ) (2025-04-25T15:39:04Z) - Understanding and Mitigating Language Confusion in LLMs [76.96033035093204]
我々は,既存の英語および多言語プロンプトを用いた15の型的多様言語の評価を行った。
Llama Instruct と Mistral のモデルでは,言語的混乱の度合いが高いことがわかった。
言語混乱は,数発のプロンプト,多言語SFT,選好調整によって部分的に緩和できることがわかった。
論文 参考訳(メタデータ) (2024-06-28T17:03:51Z) - Boosting Cross-Lingual Transfer via Self-Learning with Uncertainty
Estimation [34.97086123805344]
最近の多言語事前訓練型言語モデルは、目覚ましいゼロショット性能を実現している。
対象言語のラベルのないデータをさらに活用する自己学習フレームワークを提案する。
我々は,NER(Nond Entity Recognition)とNLI(Natural Language Inference)の2つの言語間タスクについて,40言語を網羅した不確実性で評価した。
論文 参考訳(メタデータ) (2021-09-01T05:26:46Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。