論文の概要: VIG: Visual Information Gain as a Reward Signal for Multimodal Chain-of-Thought Compression
- arxiv url: http://arxiv.org/abs/2608.21883v1
- Date: Sat, 22 Aug 2026 09:52:08 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-25 13:29:43.509331
- Title: VIG: Visual Information Gain as a Reward Signal for Multimodal Chain-of-Thought Compression
- Title(参考訳): VIG:マルチモーダル・チェーン・オブ・サート圧縮のためのリワード信号としての視覚情報ゲイン
- Authors: Wen Luo, Xiaohan Yi, Xiaotao Huang, Liqun Huang,
- Abstract要約: ビジュアル情報ゲイン(Visual Information Gain、VIG)は、画像が予測の不確実性をどれだけ減らすかによって、各推論トークンをスコアする情報理論的な報酬である。
VIGは、6つの主要なマルチモーダル推論ベンチマークと3つのQwen3-VL-Thinkingモデルサイズ間の精度-効率トレードオフを一貫して改善する。
- 参考スコア(独自算出の注目度): 6.871563715334457
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Multimodal large reasoning models often rely on long Chain-of-Thought (CoT) traces in which a substantial fraction of tokens, such as repeated visual descriptions, self-reflection, and other visually-disengaged filler, inflate inference cost without contributing to the answer. Existing CoT compression methods optimize output length but never measure whether a reasoning token is actually grounded in the image. We propose \textbf{VIG} (Visual Information Gain), an information-theoretic GRPO reward that scores each reasoning token by how much the image reduces its predictive uncertainty. VIG is computed online from two forward passes of the same policy, one with and one without the image, so no reference chains, external annotations, or auxiliary reward models are needed. Across six main multimodal reasoning benchmarks and three Qwen3-VL-Thinking model sizes (2B/4B/8B), plus an additional R1-Onevision-Bench evaluation on 8B, VIG consistently improves the accuracy--efficiency trade-off, supporting our central claim: \emph{efficient multimodal reasoning emerges from raising visual information density, where every reasoning token earns its place by anchoring to the image, rather than from imposing a length budget.} Our source code is available at https://github.com/chaser682/vig.
- Abstract(参考訳): マルチモーダルな大きな推論モデルは、しばしば長いチェーン・オブ・ソート(CoT)のトレースに依存しており、繰り返し視覚的な記述、自己反射、その他の視覚的に切り離されたフィラーなどのトークンが、答えに寄与することなく推論コストを増大させる。
既存のCoT圧縮手法は出力長を最適化するが、推論トークンが実際に画像に接地されているかどうかを計測することはない。
本稿では,情報理論のGRPO報酬である‘textbf{VIG}(Visual Information Gain)を提案する。
VIGは、同じポリシーの2つの前方パスからオンラインに計算され、1つは画像のないもので、参照チェーン、外部アノテーション、補助報酬モデルは必要ない。
6つの主要なマルチモーダル推論ベンチマークと3つのQwen3-VL-Thinkingモデルサイズ(2B/4B/8B)に加えて、R1-Onevision-Benchによる8Bの評価が追加され、VIGは一貫して精度と効率のトレードオフを改善し、我々の中心的主張をサポートする。
ソースコードはhttps://github.com/chaser682/vig.orgから入手可能です。
関連論文リスト
- Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models [54.097381112138]
現在の自己進化型LMMは、イメージキャプションや視覚的質問応答といった視覚言語理解タスクに苦慮している。
モデルの視覚条件を直接正規化する,純粋に教師なしの自己進化型フレームワークであるVISEを提案する。
VISEは、専門的な役割、外部報酬モデル、アノテーションなしで単一のモデル内で動作します。
論文 参考訳(メタデータ) (2026-06-25T17:59:55Z) - ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better [59.29940512530982]
推論プロセスに視覚的ヒントを動的に統合するフレームワークChainVを提案する。
提案手法は,特に算数集約ベンチマークにおいて,推論精度と効率を大幅に向上させる。
論文 参考訳(メタデータ) (2025-11-21T10:11:17Z) - VR-Thinker: Boosting Video Reward Models through Thinking-with-Image Reasoning [49.610569478718226]
マルチモーダル報酬モデル(RM)は、視覚生成モデルのトレーニング後を大幅に改善した。
VideoReward Thinker (VR-Thinker)は、RMに視覚的推論操作と視覚的メモリウィンドウを備えた思考とイメージのフレームワークである。
提案手法は,映像選好ベンチマークにおいて,オープンソースモデル間で最先端の精度を提供する。
論文 参考訳(メタデータ) (2025-10-12T09:29:50Z) - Reinforcing Video Reasoning with Focused Thinking [65.85683941058916]
本稿では,集中的思考と深い報酬の粒度で視覚的推論を強化する新しいフレームワークであるTW-GRPOを提案する。
具体的には,高情報密度のトークンを優先するトークン重み付け機構を用いる。
また,シングルチョイスからマルチチョイスQAタスクにシフトすることで,RLトレーニングを再構築する。
論文 参考訳(メタデータ) (2025-05-30T15:42:19Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。