論文の概要: Looking under the Wrong Lamppost: On the Limitations of Automated Translation Quality Estimation
- arxiv url: http://arxiv.org/abs/2608.03577v1
- Date: Tue, 04 Aug 2026 12:33:13 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-05 15:30:23.188562
- Title: Looking under the Wrong Lamppost: On the Limitations of Automated Translation Quality Estimation
- Title(参考訳): Wrong Lamppostに見る:自動翻訳品質推定の限界について
- Authors: Serge Gladkoff, Angelika Vaasa, Sue Ellen Wright, Ingemar Strandvik, Lifeng Han,
- Abstract要約: 我々は、現在のQEシステムは、現実世界の翻訳において信頼できるスタンドアロンツールとして機能するように構造的に不整合であると主張している。
分離されたセグメントのレベルでの翻訳の質の評価は、凝集度、コヒーレンス、およびスタイル的および修辞的テキストの特徴を欠く傾向にあるため、問題となる。
- 参考スコア(独自算出の注目度): 3.0039231903051875
- License: http://creativecommons.org/licenses/by-nc-sa/4.0/
- Abstract: Automation of Translation Quality Estimation (QE) has emerged as a widely discussed approach to managing translation quality at scale, and a growing number of tools and technologies have been released in pursuit of this goal. However, the proliferation of new QE systems has not always been accompanied by robust, transparent, and reproducible research and testing. This gap deserves critical scrutiny. This paper examines some fundamental limitations of the QE technology from both theoretical and empirical perspectives, arguing that current QE systems are structurally ill-equipped to serve as reliable standalone tools in real-world translation workflows. The reviewed evidence suggests that QE suffers from a range of interrelated and largely unresolved limitations. Most fundamentally, the evaluation of the quality of translation at the level of isolated segments is problematic because it tends to miss out on cohesion, coherence, and stylistic and rhetorical text features. In addition, empirical research documents several other limitations and flaws, including failure to generalize, systematic biases, overfitting and distribution collapse, performance gaps, error annotation challenges, and data scarcity. These are structural limitations arising from the complexity of human language and translation as a cognitive and communicative act - limitations that more data and better architectures have so far not overcome. Consequently, segment-level QE scores should not be used as a standalone basis for routing, release, or review bypass in production; we argue future work should focus on automating human evaluation grounded in MQM.
- Abstract(参考訳): 翻訳品質評価の自動化(QE)は,翻訳品質を大規模に管理する手法として広く議論されている。
しかし、新しいQEシステムの増殖は必ずしも、堅牢で透明で再現可能な研究と試験を伴ってはいない。
このギャップは批判的な監視に値する。
本稿では,QE技術の基本的限界を理論的・実証的両面から検討し,現在のQEシステムは実世界の翻訳ワークフローにおいて信頼性の高いスタンドアロンツールとして機能するように構造的に不整合であると主張した。
レビューされた証拠は、QEが様々な関連性があり、ほとんど解決されていない制限に悩まされていることを示唆している。
第一に、分離されたセグメントのレベルでの翻訳の質の評価は、凝集、コヒーレンス、および様式的および修辞的テキストの特徴を欠く傾向があるため、問題となる。
さらに、実証的研究は、一般化の失敗、系統的バイアス、過度な適合と分散の崩壊、パフォーマンスのギャップ、エラーアノテーションの問題、データ不足など、いくつかの制限と欠陥を文書化している。
これらは、人間の言語と、認知的かつコミュニケーション的な行為としての翻訳の複雑さから生じる構造的な制限です。
したがって、セグメントレベルのQEスコアは、本番環境でのルーティング、リリース、レビューバイパスのスタンドアロンベースとして使うべきではない。
関連論文リスト
- Semi-Synthetic Parallel Data for Translation Quality Estimation: A Case Study of Dataset Building for an Under-Resourced Language Pair [0.0]
本研究は、英語からヘブライ語へのQEのための半合成並列データセットを提案する。
専門的に翻訳された英語・ヘブライ語セグメントを、我々の資源から取り入れ、最高品質スコアを付与した。
言語的問題、特に性別と数字の合意に関する問題に対処するために、制御された翻訳エラーが導入された。
論文 参考訳(メタデータ) (2026-03-12T09:48:34Z) - Beyond Scalar Scores: Reinforcement Learning for Error-Aware Quality Estimation of Machine Translation [10.050982803590903]
品質評価は、参照翻訳に頼ることなく、機械翻訳(MT)出力の品質を評価することを目的としている。
重度リソース不足の言語ペアであるMalayalamに、英語のための最初のセグメントレベルQEデータセットを導入する。
ALOPE-RLは、効率的なアダプタを訓練するポリシーベースの強化学習フレームワークである。
論文 参考訳(メタデータ) (2026-02-09T12:42:41Z) - The AI Imperative: Scaling High-Quality Peer Review in Machine Learning [49.87236114682497]
AIによるピアレビューは、緊急の研究とインフラの優先事項になるべきだ、と私たちは主張する。
我々は、事実検証の強化、レビュアーのパフォーマンスの指導、品質改善における著者の支援、意思決定におけるAC支援におけるAIの具体的な役割を提案する。
論文 参考訳(メタデータ) (2025-06-09T18:37:14Z) - Do LLMs Understand Your Translations? Evaluating Paragraph-level MT with Question Answering [68.3400058037817]
本稿では,TREQA(Translation Evaluation via Question-Answering)について紹介する。
我々は,TREQAが最先端のニューラルネットワークとLLMベースのメトリクスより優れていることを示し,代用段落レベルの翻訳をランク付けする。
論文 参考訳(メタデータ) (2025-04-10T09:24:54Z) - Perturbation-based QE: An Explainable, Unsupervised Word-level Quality
Estimation Method for Blackbox Machine Translation [12.376309678270275]
摂動に基づくQEは、単に摂動入力元文上で出力されるMTシステムを分析することで機能する。
我々のアプローチは、教師付きQEよりも、翻訳における性別バイアスや単語センスの曖昧さの誤りを検出するのに優れている。
論文 参考訳(メタデータ) (2023-05-12T13:10:57Z) - Measuring Uncertainty in Translation Quality Evaluation (TQE) [62.997667081978825]
本研究は,翻訳テキストのサンプルサイズに応じて,信頼区間を精度良く推定する動機づけた研究を行う。
我々はベルヌーイ統計分布モデリング (BSDM) とモンテカルロサンプリング分析 (MCSA) の手法を適用した。
論文 参考訳(メタデータ) (2021-11-15T12:09:08Z) - Pushing the Right Buttons: Adversarial Evaluation of Quality Estimation [25.325624543852086]
本稿では,機械翻訳(MT)システムにおける品質推定の逆検定法を提案する。
近年のSOTAによる人的判断と高い相関があるにもかかわらず、ある種の意味エラーはQEが検出する上で問題である。
第二に、平均的に、あるモデルが意味保存と意味調整の摂動を区別する能力は、その全体的な性能を予測できることが示される。
論文 参考訳(メタデータ) (2021-09-22T17:32:18Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。