論文の概要: Rethinking CD: A Reproducibility Study and Extension on the Ineffectiveness of Contrastive Decoding at Mitigating Object Hallucinations in MLLMs
- arxiv url: http://arxiv.org/abs/2607.25196v1
- Date: Tue, 28 Jul 2026 01:55:58 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-29 20:50:42.677091
- Title: Rethinking CD: A Reproducibility Study and Extension on the Ineffectiveness of Contrastive Decoding at Mitigating Object Hallucinations in MLLMs
- Title(参考訳): CDの再考 : MLLMにおける物体幻覚の緩和におけるコントラスト復号の非効率性に関する再現性研究と拡張
- Authors: Arnav Bendre, Guneesh Gupta, Kavish Grover, Chayan Aggarwal, Shreyansh Modi,
- Abstract要約: 大規模言語モデル(MLLM)における物体幻覚を緩和するための訓練不要戦略として、コントラスト復号法(CD)が提案されている。
本研究は,CDからの明らかな改善がしばしば刺激的であり,幻覚を減少させるため,常に強い視覚的基盤に変換されないことを示す。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Contrastive decoding (CD) has been proposed as a training-free strategy for mitigating object hallucinations in multimodal large language models (MLLMs), with reported gains on benchmarks such as POPE. However, recent work has questioned whether these gains reflect genuine improvements in visual grounding. In this study, we reproduce and extend the findings of "The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs." Specifically, we test the claim that CD induces a unidirectional output distribution shift in discriminative datasets and examine its generalizability across datasets. We also verify that the adaptive plausibility constraint (APC) reduces sampling to greedy search on both discriminative and generative benchmarks. Beyond reproduction, we rigorously study the effects of CD across generative and discriminative datasets. We conduct several experiments that provide additional insights: we analyze the logit distributions induced by different CD strategies on generative datasets, propose a proxy method and compare its performance against CD techniques, and investigate how hallucination signals propagate through each layer of the expert and amateur models. Experimental results across MME, POPE, and CHAIR using LLaVA and Qwen validate the original claims and show that the apparent improvements from CD are often spurious and do not consistently translate into stronger visual grounding for reducing hallucinations. These findings challenge the effectiveness of current contrastive decoding strategies and motivate the development of more reliable approaches for mitigating hallucinations in MLLMs.
- Abstract(参考訳): マルチモーダルな大言語モデル(MLLM)におけるオブジェクト幻覚を緩和するための訓練不要戦略として、CD(Contrastive Decoding)が提案されている。
しかし、近年の研究は、これらの成果が視覚的接地における真の改善を反映しているかどうかを疑問視している。
本研究では,「性能向上の鏡 : なぜコントラストデコードがMLLMにおける物体幻覚の緩和に失敗したのか」という知見を再現し,拡張する。
具体的には,CDが識別データセットにおける一方向の出力分布シフトを誘導するという主張を検証し,そのデータセット間の一般化可能性を検討する。
また,適応可視性制約 (APC) は, 識別的および生成的ベンチマークの両方において, グリージー検索のサンプリングを減少させることを示す。
再生を超えて、生成的および識別的データセット間でCDの効果を厳格に研究する。
生成データセット上で異なるCD戦略によって誘導されるロジット分布を分析し、プロキシ手法を提案し、その性能をCD技術と比較し、専門家およびアマチュアモデルの各層で幻覚信号がどのように伝播するかを検討する。
LLaVAとQwenを用いたMME,POPE,CHAIRにまたがる実験結果から,CDからの明らかな改善がしばしば刺激的であり,幻覚を減少させるために常に強い視覚的基盤に変換されないことが示された。
これらの知見は、現在のコントラスト復号戦略の有効性に挑戦し、MLLMにおける幻覚を緩和するためのより信頼性の高いアプローチの開発を動機付けている。
関連論文リスト
- Locate-then-Sparsify: Attribution Guided Sparse Strategy for Visual Hallucination Mitigation [68.41785694664011]
機能ステアリングのためのLate-Then-Sparsify(LTS-FS)と呼ばれるプラグアンドプレイフレームワークを提案する。
各層の幻覚関係に応じて操舵強度を制御する。
我々の枠組みは、強い性能を維持しながら幻覚を効果的に緩和する。
論文 参考訳(メタデータ) (2026-03-17T09:16:50Z) - Seeing Through the Chain: Mitigate Hallucination in Multimodal Reasoning Models via CoT Compression and Contrastive Preference Optimization [78.94590726578014]
マルチモーダル推論モデル (Multimodal reasoning model, MLRM) は幻覚の傾向が強く, 効果的な解はいまだ未発見のままである。
textbfCompression と textbfPreference textbfOptimization を組み合わせたトレーニングベースの緩和フレームワーク C3PO を提案する。
論文 参考訳(メタデータ) (2026-02-03T11:00:55Z) - ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM [16.694799255671914]
マルチモーダル大言語モデル(MLLM)は、しばしば刺激的な視覚的手がかりに過剰なコミットによって幻覚する。
本稿では,アテンション・ステアブル・コントラスト・デコーディング(ASCD)を提案する。
論文 参考訳(メタデータ) (2025-06-17T17:58:11Z) - The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs? [51.476751660784174]
そこで我々は,一連の素早い改善手法を導入し,その性能を対照的な復号化技術に対して評価する。
実験結果から, 対照的な復号化における性能向上は, 幻覚の緩和という目的とは無関係であることが判明した。
論文 参考訳(メタデータ) (2025-04-14T09:25:37Z) - Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding [25.489832294197797]
本稿では,LVLM推論における幻覚の低減を目的とした,命令コントラストデコーディング(ICD)手法を提案する。
本手法は,マルチモーダル核融合モジュールにおいて,外乱指示が幻覚を著しく悪化させるという観察に着想を得たものである。
論文 参考訳(メタデータ) (2024-03-27T16:04:47Z) - Debiasing Multimodal Large Language Models via Penalization of Language Priors [38.97645845493758]
MLLM(Multimodal Large Language Models)は、コンピュータビジョンや自然言語処理において欠かせないツールとなっている。
生成されたコンテンツは、入力画像よりも、基礎となるLarge Language Models (LLMs) の本質的な先行性によって駆動されることが多い。
本稿では、これらのバイアスを補正し、視覚情報に対するモデルの焦点をリダイレクトするための、単純でトレーニングのない2つの戦略を提案する。
論文 参考訳(メタデータ) (2024-03-08T12:35:07Z) - Alleviating Hallucinations of Large Language Models through Induced
Hallucinations [67.35512483340837]
大規模言語モデル(LLM)は、不正確な情報や製造された情報を含む応答を生成するために観察されている。
幻覚を緩和するための単純なtextitInduce-then-Contrast Decoding (ICD) 戦略を提案する。
論文 参考訳(メタデータ) (2023-12-25T12:32:49Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。