論文の概要: Understanding Evaluation Illusion in Diffusion Large Language Models
- arxiv url: http://arxiv.org/abs/2606.29228v2
- Date: Wed, 01 Jul 2026 01:03:45 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-02 15:15:53.060635
- Title: Understanding Evaluation Illusion in Diffusion Large Language Models
- Title(参考訳): 拡散大言語モデルにおける評価イリュージョンの理解
- Authors: Hengxiang Zhang, Jiaxi Ren, Renchunzi Xie, Hongxin Wei,
- Abstract要約: 拡散大言語モデル(dLLM)は、生成品質を維持するために多くのデノベーションステップを必要とする。
既往の研究では, 同一視される評価条件下においても, 不整合評価結果が報告されている。
本稿では,dLLMにおける復号化手法の信頼性評価のための実践的ガイドラインを提案する。
- 参考スコア(独自算出の注目度): 19.57955271138274
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Despite the capability of parallel decoding, diffusion large language models (dLLMs) require many denoising steps to maintain generation quality, motivating recent research on efficient decoding strategies. However, existing studies have reported inconsistent evaluation results even under seemingly identical evaluation settings, risking biased conclusions about dLLM decoding methods. To understand this evaluation concern, we conduct a rigorous evaluation of current decoding methods for dLLMs across diverse evaluation settings. Surprisingly, our analysis reveals that the ranking of decoding methods is highly sensitive to the choice of prompt templates. Single-template evaluation can lead to an illusion that decoding methods improve inference efficiency without performance degradation. Through comprehensive experiments, we find that current parallel decoding methods consistently underperform the single-token decoding baseline, failing to overcome the speed-quality trade-off. We further identify this evaluation inconsistency as the high sensitivity of parallel decoding methods to minor variations in prompt templates. Our experiments show that an effective prompt template can achieve strong evaluation results even with fewer denoising steps, markedly outperforming the marginal gain from increasing denoising steps. Beyond prompt templates, our experiments indicate that overlooked evaluation settings can also notably affect the assessment of decoding methods. Based on these findings, we propose practical guidelines for the reliable evaluation of decoding methods in dLLMs.
- Abstract(参考訳): 並列復号化の能力にもかかわらず、拡散大言語モデル(dLLM)は生成品質を維持するために多くの復号化ステップを必要としており、最近の効率的な復号化戦略の研究を動機付けている。
しかし, 既存の研究では, dLLM復号法に関する偏りのある結論をリスクとして, 一見同一の評価条件下においても不整合性評価結果が報告されている。
この評価問題を理解するため、我々は様々な評価設定において、dLLMの現在の復号法を厳格に評価する。
解析の結果,デコード手法のランク付けは,プロンプトテンプレートの選択に非常に敏感であることが判明した。
単一テンプレート評価は、デコード手法が性能劣化を伴わずに推論効率を向上させるという錯覚をもたらす可能性がある。
包括的実験により、現在の並列復号法はシングルトーケンデコードベースラインを一貫して過小評価し、速度品質のトレードオフを克服できないことがわかった。
さらに、この評価の不整合性を、プロンプトテンプレートの小さなバリエーションに対する並列復号法の高感度性として認識する。
実効的なプロンプトテンプレートは,デノナイジングステップが少なくても高い評価結果が得られることを示し,デノナイジングステップの増加による限界ゲインを著しく上回ることを示した。
プロンプトテンプレート以外にも,見過ごされた評価設定がデコード手法の評価に顕著に影響を及ぼす可能性が示唆された。
そこで本研究では,dLLMにおける復号化手法の信頼性評価のための実践的ガイドラインを提案する。
関連論文リスト
- Unlocking LLM Code Correction with Iterative Feedback Loops [5.504955093712013]
本研究では、コード障害の評価、修正パターンの分析、推論と非推論モデルの有効性の比較を行うメトリクスを紹介する。
その結果、推論モデルは反復よりも一貫して改善され、フィードバックを活用する際に非推論モデルよりも大幅に優れています。
論文 参考訳(メタデータ) (2026-06-16T04:47:42Z) - How Efficient Are Diffusion Language Models? A Critical Examination of Efficiency Evaluation Practices [81.85465545346266]
拡散言語モデル(DLM)は、長期支配的自己回帰(AR)パラダイムに代わる有望な代替として登場した。
しかし、現在のオープンソースのDLMは、しばしばARの速度よりも優れており、現実のユーティリティを制限している。
本研究はDLMの効率に関する系統的研究であり, 先行評価手法の問題点を同定する。
論文 参考訳(メタデータ) (2025-10-21T10:00:32Z) - Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation [6.4212082894269535]
既存のリーク検出技術である置換法とn-gram法を比較した。
解析の結果,n-gram法は高いF1スコアが得られることがわかった。
MMLUとHellaSwagのクリーンバージョンを作成し、複数のLLMを再評価する。
論文 参考訳(メタデータ) (2025-05-30T06:37:39Z) - Cracking the Code: Evaluating Zero-Shot Prompting Methods for Providing Programming Feedback [2.309018557701645]
このケーススタディでは、異なるゼロショットプロンプトエンジニアリング手法を評価するための評価フレームワークを導入している。
提案手法を体系的に変更し,提案したRのプログラムエラーに対するフィードバックを分析した。
論文 参考訳(メタデータ) (2024-12-20T09:24:50Z) - A Thorough Examination of Decoding Methods in the Era of LLMs [72.65956436513241]
復号法は、次世代の予測器から実用的なタスク解決器に言語モデルを変換する上で、必須の役割を果たす。
本稿では,大規模言語モデルの文脈における様々な復号法を包括的かつ多面的に分析する。
その結果,復号法の性能は特にタスク依存的であり,アライメント,モデルサイズ,量子化などの要因に影響されていることが明らかとなった。
論文 参考訳(メタデータ) (2024-02-10T11:14:53Z) - Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding [46.485363806259265]
投機的デコーディングは、LLM(Large Language Models)推論のための新しいデコーディングパラダイムとして登場した。
復号処理の各ステップにおいて、この手法はまず、複数の将来のトークンを効率的にドラフトし、それらを並列に検証する。
本稿では,この有望な復号化パラダイムの概観と解析について述べる。
論文 参考訳(メタデータ) (2024-01-15T17:26:50Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。