論文の概要: Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs
- arxiv url: http://arxiv.org/abs/2603.27494v1
- Date: Sun, 29 Mar 2026 03:18:57 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-03-31 23:18:44.985772
- Title: Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs
- Title(参考訳): 情報ギャップとMLLMの接地損失を考慮した強化学習フレームワーク
- Authors: Xuanpu Zhao, Zhentao Tan, Dianmo Sheng, Tianxiang Chen, Yao Liu, Yue Wu, Tao Gong, Qi Chu, Nenghai Yu,
- Abstract要約: MLLMは画像トリミングツールを自律的に利用し、質問応答のための関心領域を分析する。
我々は,このモデルが大域的な入力に強く依存していることと,収穫領域の細部への弱い依存を実証する。
監視を必要としない新しい2段階強化学習フレームワークを提案する。
- 参考スコア(独自算出の注目度): 58.32326205291745
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: To enhance the perception and reasoning capabilities of multimodal large language models in complex visual scenes, recent research has introduced agent-based workflows. In these works, MLLMs autonomously utilize image cropping tool to analyze regions of interest for question answering. While existing training strategies, such as those employing supervised fine-tuning and reinforcement learning, have made significant progress, our empirical analysis reveals a key limitation. We demonstrate the model's strong reliance on global input and its weak dependence on the details within the cropped region. To address this issue, we propose a novel two-stage reinforcement learning framework that does not require trajectory supervision. In the first stage, we introduce the ``Information Gap" mechanism by adjusting the granularity of the global image. This mechanism trains the model to answer questions by focusing on cropped key regions, driven by the information gain these regions provide. The second stage further enhances cropping precision by incorporating a grounding loss, using a small number of bounding box annotations. Experiments show that our method significantly enhances the model's attention to cropped regions, enabling it to achieve state-of-the-art performance on high-resolution visual question-answering benchmarks. Our method provides a more efficient approach for perceiving and reasoning fine-grained details in MLLMs. Code is available at: https://github.com/XuanPu-Z/LFPC.
- Abstract(参考訳): 複雑な視覚シーンにおけるマルチモーダルな大言語モデルの知覚と推論能力を高めるため、近年ではエージェントベースのワークフローを導入している。
これらの研究において、MLLMは画像トリミングツールを自律的に利用し、質問応答のための関心領域を分析する。
教師付き微調整や強化学習などの既存の訓練戦略は大きな進歩を遂げているが、実証分析では重要な限界が明らかになっている。
我々は,このモデルが大域的な入力に強く依存していることと,収穫領域の細部への弱い依存を実証する。
この問題に対処するために,軌道監視を必要としない新しい2段階強化学習フレームワークを提案する。
第1段階では,グローバル画像の粒度を調整して「情報ギャップ」機構を導入する。
このメカニズムは、これらの領域が提供する情報によって駆動される、収穫されたキー領域に焦点を当てて、モデルに質問に答えるよう訓練する。
第2段階はさらに、少数のバウンディングボックスアノテーションを使用して、接地損失を組み込むことで、収穫精度を向上する。
実験の結果,本手法は収穫地への関心を著しく高め,高精細度視覚質問応答ベンチマークにおける最先端性能を実現することができることがわかった。
本手法は, MLLMの細粒度を知覚し, 推論する上で, より効率的な手法を提供する。
コードは、https://github.com/XuanPu-Z/LFPC.comで入手できる。
関連論文リスト
- FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning [62.11389260206383]
textscFineRSは、非常に小さなオブジェクトをセグメント化するための2段階のMLLMベースの強化学習フレームワークである。
textscFineRS-4kは,属性レベルの推論に基づくMLLMの評価と,微妙で小規模なターゲットに対する画素レベルのセグメンテーションのための新しいデータセットである。
論文 参考訳(メタデータ) (2025-10-24T10:14:17Z) - CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models [16.91226496250909]
マルチモーダルな理解は、粗いものから細かいものへと、2つの段階に分けられる。
第1段階では,MLLMに回答のほぼ面積を特定するよう促す。
第2段階では、視覚的なプロンプトエンジニアリングにより、関連する領域に対するモデルの焦点をさらに強化する。
論文 参考訳(メタデータ) (2024-12-22T05:42:40Z) - USDRL: Unified Skeleton-Based Dense Representation Learning with Multi-Grained Feature Decorrelation [24.90512145836643]
本稿では,特徴デコレーションに基づく統一骨格に基づくDense Representation Learningフレームワークを提案する。
我々のアプローチは現在のSOTA(State-of-the-art)アプローチよりも大幅に優れています。
論文 参考訳(メタデータ) (2024-12-12T12:20:27Z) - Recognize Any Regions [55.76437190434433]
RegionSpotは、ローカライゼーション基盤モデルから位置認識ローカライゼーション知識と、ViLモデルからのセマンティック情報を統合する。
オープンワールドオブジェクト認識の実験では、私たちのRereaSpotは、以前の代替よりも大きなパフォーマンス向上を実現しています。
論文 参考訳(メタデータ) (2023-11-02T16:31:49Z) - USER: Unified Semantic Enhancement with Momentum Contrast for Image-Text
Retrieval [115.28586222748478]
Image-Text Retrieval (ITR) は、与えられたクエリに意味のあるターゲットインスタンスを、他のモダリティから検索することを目的としている。
既存のアプローチは通常、2つの大きな制限に悩まされる。
論文 参考訳(メタデータ) (2023-01-17T12:42:58Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。