Fugu-MT 論文翻訳(概要): MPCAR: Multi-Perspective Contextual Augmentation for Enhanced Visual Reasoning in Large Vision-Language Models

論文の概要: MPCAR: Multi-Perspective Contextual Augmentation for Enhanced Visual Reasoning in Large Vision-Language Models

arxiv url: http://arxiv.org/abs/2508.12400v1
Date: Sun, 17 Aug 2025 15:25:01 GMT
ステータス: 翻訳完了
システム内更新日: 2025-08-19 14:49:10.744931
Title: MPCAR: Multi-Perspective Contextual Augmentation for Enhanced Visual Reasoning in Large Vision-Language Models
Title（参考訳）: MPCAR:大規模視覚言語モデルにおける拡張視覚推論のための多視点文脈拡張
Authors: Amirul Rahman, Qiang Xu, Xueying Huang,
Abstract要約: Multi-Perspective Contextual Augmentation for Reasoning (MPCAR)は、LVLM(Large Vision-Language Models)を強化するために設計された新しい推論時間戦略である。第一に、LVLMは様々な角度から N の多様で相補的な記述や予備的推論経路を生成し、第二に、これらの記述は、元の質問とインテリジェントに統合され、包括的な文脈拡張プロンプトを構築し、最後に、このリッチ化されたプロンプトは、深い推論と最終回答生成のために究極の LVLM を導く。
参考スコア（独自算出の注目度）: 7.702194892874595
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Abstract: Despite significant advancements, Large Vision-Language Models (LVLMs) continue to face challenges in complex visual reasoning tasks that demand deep contextual understanding, multi-angle analysis, or meticulous detail recognition. Existing approaches often rely on single-shot image encoding and prompts, limiting their ability to fully capture nuanced visual information. Inspired by the notion that strategically generated "additional" information can serve as beneficial contextual augmentation, we propose Multi-Perspective Contextual Augmentation for Reasoning (MPCAR), a novel inference-time strategy designed to enhance LVLM performance. MPCAR operates in three stages: first, an LVLM generates N diverse and complementary descriptions or preliminary reasoning paths from various angles; second, these descriptions are intelligently integrated with the original question to construct a comprehensive context-augmented prompt; and finally, this enriched prompt guides the ultimate LVLM for deep reasoning and final answer generation. Crucially, MPCAR achieves these enhancements without requiring any fine-tuning of the underlying LVLM's parameters. Extensive experiments on challenging Visual Question Answering (VQA) datasets, including GQA, VQA-CP v2, and ScienceQA (Image-VQA), demonstrate that MPCAR consistently outperforms established baseline methods. Our quantitative results show significant accuracy gains, particularly on tasks requiring robust contextual understanding, while human evaluations confirm improved coherence and completeness of the generated answers. Ablation studies further highlight the importance of diverse prompt templates and the number of generated perspectives. This work underscores the efficacy of leveraging LVLMs' inherent generative capabilities to enrich input contexts, thereby unlocking their latent reasoning potential for complex multimodal tasks.
Abstract（参考訳）: 大幅な進歩にもかかわらず、LVLM(Large Vision-Language Models)は、深い文脈理解、多角分析、細部認識を必要とする複雑な視覚的推論タスクの課題に直面し続けている。既存のアプローチはしばしばシングルショット画像エンコーディングとプロンプトに依存しており、ニュアンス付き視覚情報をフルにキャプチャする能力を制限している。戦略的に生成された「付加的」情報が有益な文脈拡張に役立つという概念に着想を得て,LVLMの性能向上を目的とした新しい推論時戦略MPCARを提案する。第一に、LVLMは様々な角度から N の多様かつ相補的な記述や予備的推論経路を生成し、第二に、これらの記述は、元の質問とインテリジェントに統合され、包括的な文脈拡張プロンプトを構築し、最後に、このリッチ化されたプロンプトは、深い推論と最終回答生成のために究極の LVLM を導く。重要なことに、MPCARは、基礎となるLVLMのパラメータを微調整することなく、これらの拡張を実現している。 GQA、VQA-CP v2、ScienceQA(Image-VQA)など、VQA(Visual Question Answering)データセットへの挑戦に関する大規模な実験は、MPCARが確立されたベースラインメソッドを一貫して上回っていることを実証している。人間の評価は, 結果の一貫性と完全性の向上を裏付ける一方で, 特に, 文脈理解の堅牢性を必要とするタスクにおいて, 有意な精度向上を示す。アブレーション研究は、多様なプロンプトテンプレートの重要性と生成された視点の数をさらに強調している。この研究は、LVLMs固有の生成能力を活用して入力コンテキストを豊かにすることで、複雑なマルチモーダルタスクに対する潜在的推論能力を解放する効果を裏付けるものである。

論文の概要: MPCAR: Multi-Perspective Contextual Augmentation for Enhanced Visual Reasoning in Large Vision-Language Models

関連論文リスト