論文の概要: Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict
- arxiv url: http://arxiv.org/abs/2609.00550v1
- Date: Tue, 01 Sep 2026 01:37:18 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-02 16:31:36.197778
- Title: Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict
- Title(参考訳): 同じセマンティックスと異なる結果:知識紛争下におけるマルチモーダルLLMのモダリティロバスト性について
- Authors: Jungyeon Lee, Yejin Yoon, Taeuk Kim,
- Abstract要約: MLLM(Multimodal large language model)は、異種形式の文脈的証拠としてますます多く提供される。
モデルがパラメトリックな知識と矛盾する場合、これらの曲面がいかに一貫して処理されるかは、いまだに不明である。
13のMLLMと2つのデータセットにまたがる知識衝突下でのモダリティの堅牢性について検討し、それらがロバストから遠く離れていることを発見した。
- 参考スコア(独自算出の注目度): 8.080331183351534
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Multimodal large language models (MLLMs) are increasingly provided with contextual evidence in heterogeneous forms: as a text passage, as a rendered image of the same passage, or as both together. However, it remains unclear how consistently these surface forms are processed, especially when the evidence conflicts with the model's parametric knowledge. We study modality robustness under knowledge conflict across 13 MLLMs and two datasets, and find them far from robust. (1) Contrary to common belief, models favor a context that contradicts parametric knowledge more readily in image form than in text form; (2) when a contradicting text and image are presented together, the preferred modality is essentially arbitrary, varying with input order, model, and dataset. We further demonstrate that this instability has practical consequences: it degrades performance in multimodal RAG and can be exploited by adversarial attacks. To alleviate this brittleness, we examine several simple techniques---prompting, steering, supervised fine-tuning (SFT), and direct preference optimization; the majority prove ineffective, whereas SFT achieves moderate success. We therefore call for greater awareness of this inconsistency and argue that it is fundamental, demanding attention at multiple training stages.
- Abstract(参考訳): MLLM(Multimodal large language model)は、テキストパスとして、同じパスのレンダリングイメージとして、あるいは両方で、異種形式の文脈的エビデンスを提供するようになっている。
しかしながら、これらの曲面がいかに一貫して処理されるかは、特に証拠がモデルのパラメトリック知識と矛盾する場合に明らかでない。
13のMLLMと2つのデータセットにまたがる知識衝突下でのモダリティロバスト性について検討し、それらがロバストから遠く離れていることを発見した。
1) 一般的な信念とは対照的に、モデルは、テキスト形式よりも画像形式でパラメトリック知識と矛盾する文脈を好んでおり、(2)矛盾するテキストと画像が一緒に提示される場合、好ましくは任意であり、入力順序、モデル、データセットによって変化する。
さらに、この不安定性は、マルチモーダルRAGの性能を低下させ、敵攻撃によって悪用できるという、実用的な結果をもたらすことを実証する。
この脆さを緩和するために, プロンプティング, ステアリング, 教師付き微調整(SFT), 直接選好最適化などいくつかの簡単な手法を検討した。
したがって、我々は、この矛盾に対するより深い認識を呼び、それは基本的なものであり、複数の訓練段階において注意が必要であると論じる。
関連論文リスト
- Analyzing Reasoning Consistency in Large Multimodal Models under Cross-Modal Conflicts [74.47786985522762]
テキスト慣性(textual inertia)と呼ばれる重要な障害モードを特定し、矛盾する視覚的証拠を無視しながら、モデルは間違ったテキストに盲目的に固執する傾向がある。
本稿では,多種多様なLMMの推論連鎖に摂動を構造的に注入するLogicGraph摂動プロトコルを提案する。
その結果,10%未満の症例で自己修正が成功し,主に視覚的テキスト誤りの伝播に寄与することが判明した。
論文 参考訳(メタデータ) (2026-01-07T16:39:34Z) - When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models [10.106066580331584]
我々は,画像,ビデオ,オーディオ,時系列,グラフなど多種多様なデータモダリティにまたがるテキスト優位性を,初めて体系的に調査した。
奥行き分析では,非テクスチュアルなモダリティにおける高度トークン冗長性からの注意の希釈,融合アーキテクチャ設計の影響,テキスト入力を暗黙的に好むタスクの定式化という,3つの根本原因を明らかにした。
論文 参考訳(メタデータ) (2025-08-14T11:44:52Z) - Continual Multimodal Contrastive Learning [99.53621521696051]
MCL(Multimodal Contrastive Learning)は、異なるモダリティを整列し、関節空間におけるマルチモーダル表現を生成する。
マルチモーダルデータは単一のプロセスで収集されることはめったになく、スクラッチからのトレーニングは計算コストがかかる。
本稿では, 安定性と塑性の2つの原理によりCMCLを定式化する。
理論的には、二辺から部分空間への勾配の更新を計画する、新しい最適化に基づく手法を導出する。
論文 参考訳(メタデータ) (2025-03-19T07:57:08Z) - Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models [33.76903352835436]
LVLM(Large Vision-Language Models)は、マルチモーダル入力をキャプチャし、推論する能力を示す。
これらのモデルは、そのビジョンと言語コンポーネント間の表現された知識の不整合から生じるパラメトリックな知識の衝突を招きやすい。
我々は、それらを検出し、解釈し、緩和するための体系的なアプローチを提案する。
論文 参考訳(メタデータ) (2024-10-04T17:59:28Z) - Cross-Attention is Not Enough: Incongruity-Aware Dynamic Hierarchical
Fusion for Multimodal Affect Recognition [69.32305810128994]
モダリティ間の同調性は、特に認知に影響を及ぼすマルチモーダル融合の課題となる。
本稿では,動的モダリティゲーティング(HCT-DMG)を用いた階層型クロスモーダルトランスを提案する。
HCT-DMG: 1) 従来のマルチモーダルモデルを約0.8Mパラメータで上回り、2) 不整合が認識に影響を及ぼすハードサンプルを認識し、3) 潜在レベルの非整合性をクロスモーダルアテンションで緩和する。
論文 参考訳(メタデータ) (2023-05-23T01:24:15Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。