論文の概要: Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap
- arxiv url: http://arxiv.org/abs/2604.16256v1
- Date: Fri, 17 Apr 2026 17:15:18 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-04-20 22:00:20.021724
- Title: Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap
- Title(参考訳): 視覚-言語モデルによる視覚推論は真に達成されるか? : モダリティギャップの厳密な研究
- Authors: Yige Xu, Yongjie Wang, Zizhuo Wu, Kaisong Song, Jun Lin, Zhiqi Shen,
- Abstract要約: 近年,視覚言語モデル (VLM) の推論が注目されている。
我々はクロスモーダル比較の制御のために設計された新しいマルチモーダル推論ベンチマークであるCrossMathを紹介する。
我々は,同じタスク関連情報を保証するために,テキストのみ,画像のみ,画像+テキストフォーマットで各問題を構築する。
- 参考スコア(独自算出の注目度): 14.392019191368746
- License: http://creativecommons.org/licenses/by-nc-nd/4.0/
- Abstract: Reasoning in vision-language models (VLMs) has recently attracted significant attention due to its broad applicability across diverse downstream tasks. However, it remains unclear whether the superior performance of VLMs stems from genuine vision-grounded reasoning or relies predominantly on the reasoning capabilities of their textual backbones. To systematically measure this, we introduce CrossMath, a novel multimodal reasoning benchmark designed for controlled cross-modal comparisons. Specifically, we construct each problem in text-only, image-only, and image+text formats guaranteeing identical task-relevant information, verified by human annotators. This rigorous alignment effectively isolates modality-specific reasoning differences while eliminating confounding factors such as information mismatch. Extensive evaluation of state-of-the-art VLMs reveals a consistent phenomenon: a substantial performance gap between textual and visual reasoning. Notably, VLMs excel with text-only inputs, whereas incorporating visual data (image+text) frequently degrades performance compared to the text-only baseline. These findings indicate that current VLMs conduct reasoning primarily in the textual space, with limited genuine reliance on visual evidence. To mitigate this limitation, we curate a CrossMath training set for VLM fine-tuning. Empirical evaluations demonstrate that fine-tuning on this training set significantly boosts reasoning performance across all individual and joint modalities, while yielding robust gains on two general visual reasoning tasks. Source code is available at https://github.com/xuyige/CrossMath.
- Abstract(参考訳): 視覚言語モデル(VLM)の推論は、様々な下流タスクにまたがる広範囲な適用性により、近年大きな注目を集めている。
しかし、VLMの優れた性能が真の視覚的根拠による推論に由来するのか、あるいはテキストバックボーンの推論能力に大きく依存しているかは、まだ不明である。
これを体系的に測定するために、クロスモーダル比較を制御するために設計された新しいマルチモーダル推論ベンチマークであるクロスマス(CrossMath)を導入する。
具体的には、人間のアノテータが検証したタスク関連情報を同一に保証する、テキストのみ、画像のみ、画像+テキストフォーマットで各問題を構築する。
この厳密なアライメントは、情報ミスマッチのような相反する要因を排除しつつ、モダリティ固有の推論の違いを効果的に分離する。
最先端のVLMを広範囲に評価すると、一貫した現象が明らかになる。
特に、VLMはテキストのみの入力に優れ、ビジュアルデータ(画像+テキスト)はテキストのみのベースラインに比べて性能が劣化する。
これらの結果は、現在のVLMが主にテキスト空間で推論を行い、視覚的証拠に依存していることを示している。
この制限を緩和するため、VLMファインチューニングのためのクロスマストレーニングセットをキュレートする。
実験的な評価では、このトレーニングセットの微調整は、すべての個人および共同モダリティにおける推論性能を著しく向上させ、一方、2つの一般的な視覚的推論タスクに対して頑健な利得をもたらすことが示されている。
ソースコードはhttps://github.com/xuyige/CrossMath.comで入手できる。
関連論文リスト
- Perception-Aware Multimodal Spatial Reasoning from Monocular Images [57.42071289037214]
単眼画像からの空間的推論は 自律運転には不可欠です
現在のヴィジュアルランゲージモデル(VLM)は、微粒な幾何学的知覚に苦慮している。
本稿では,VLMを明示的な対象中心の接地能力を持つ知覚認識型マルチモーダル推論フレームワークを提案する。
論文 参考訳(メタデータ) (2026-03-07T02:05:12Z) - Can Unified Generation and Understanding Models Maintain Semantic Equivalence Across Different Output Modalities? [61.533560295383786]
Unified Multimodal Large Language Models (U-MLLM) は、単一のアーキテクチャ内で理解と生成を統合する。
我々は,U-MLLMが画像のモダリティにおいて同じ結果をレンダリングするために必要な場合,意味的等価性を維持することができないことを観察する。
VGUBenchは、推論ロジックを生成の忠実性から切り離すためのフレームワークである。
論文 参考訳(メタデータ) (2026-02-27T06:23:56Z) - Do Images Speak Louder than Words? Investigating the Effect of Textual Misinformation in VLMs [17.56537230934894]
Vision-Language Models (VLM)は、Visual-Question-Answering (VQA)ベンチマークで強力なマルチモーダル推論能力を示している。
これらのモデルが、誤解を招くテキストのプロンプトに対して脆弱であることを示し、しばしば矛盾するテキストを支持する明確な視覚的証拠をオーバーライドしている。
論文 参考訳(メタデータ) (2026-01-27T05:04:38Z) - Rethinking the Text-Vision Reasoning Imbalance in MLLMs through the Lens of Training Recipes [54.374410871041164]
MLLM(Multimodal large language model)は、視覚・言語タスクにおいて強力な機能を示す。
近年の研究では、視覚的・テキスト的モダリティ間の推論能力の不均衡が指摘されている。
我々は、この現象を、テキスト中心と視覚中心の入力のパフォーマンス格差として定義される、テクティモダリティギャップと呼ぶ。
論文 参考訳(メタデータ) (2025-10-26T21:06:13Z) - Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT Reasoning [53.790502697674754]
本稿では、画像入力を重要な推論段階に移行する戦略であるTake-Allong Visual Conditioning (TVC)を提案する。
TVCは、推論を通して視覚的なコンポーネントへの注意を維持するのに役立つ。
提案手法は,5つの数学的推論ベンチマークにおいて,最先端の性能を平均で達成する。
論文 参考訳(メタデータ) (2025-03-17T16:45:12Z) - Words or Vision: Do Vision-Language Models Have Blind Faith in Text? [34.88114876390461]
VLM(Vision-Language Models)は、視覚中心のタスクに対する視覚情報とテキスト情報の統合に優れる。
視覚中心設定における視覚データや様々なテキスト入力に直面するVLMのモダリティ嗜好について検討する。
不整合が発生した場合、VLMは視覚的データよりもテキストデータを不均等に信頼する。
論文 参考訳(メタデータ) (2025-03-04T02:21:07Z) - Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios [69.00444996464662]
RIV-CoT(Retrieval-based Interleaved Visual Chain-of-Thought法)を提案する。
実験の結果, RIV-CoTの解答精度は3.1%向上し, バニラCoTの解答精度は4.6%向上した。
論文 参考訳(メタデータ) (2025-01-08T18:31:16Z) - Multimodal Causal Reasoning Benchmark: Challenging Vision Large Language Models to Discern Causal Links Across Modalities [19.923665989164387]
MuCRはMultimodal Causal Reasoningベンチマークであり、合成シアム画像とテキストペアを利用してMLLMに挑戦する。
実験の結果,現在のMLLMはテキスト環境下での性能に比べ,マルチモーダル因果推論では不足していることがわかった。
本稿では,視覚的手がかりをより強調するVcCoT戦略を提案し,その効果がマルチモーダル因果推論の強化に有効であることを確認した。
論文 参考訳(メタデータ) (2024-08-15T12:04:32Z) - Improving Visual Commonsense in Language Models via Multiple Image Generation [41.565399860320966]
既存の大規模言語モデル(LLM)は、主にテキストデータのみを使用して訓練されている。
視覚言語モデルは視覚的に指向するタスクに優れており、基本的なコモンセンス推論のような視覚的でないタスクでは失敗することが多い。
この分散は、基本的なテキストベースの言語推論と堅牢な視覚的理解の統合という、重要な課題を浮き彫りにする。
論文 参考訳(メタデータ) (2024-06-19T15:17:10Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。