論文の概要: Mitigating the Reasoning Tax in Vision-Language Fine-Tuning with Input-Adaptive Depth Aggregation
- arxiv url: http://arxiv.org/abs/2603.26330v1
- Date: Fri, 27 Mar 2026 11:47:39 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-03-30 21:49:48.478906
- Title: Mitigating the Reasoning Tax in Vision-Language Fine-Tuning with Input-Adaptive Depth Aggregation
- Title(参考訳): 入力適応深さアグリゲーションを用いた視覚言語ファインチューニングにおける推論税の緩和
- Authors: Yiming Ren, Yujiu Yang, Junjie Wang,
- Abstract要約: 視覚的インストラクションデータに対するSFT(Supervised Fine-tuning)は、しばしば視覚言語モデル(VLM)の知覚能力を向上し、推論性能を低下させる。
この劣化が深度表現の障害的アクセスと関係しているかどうかを考察し、固定された深度集合でさえ推論を著しく復元することを示した。
IADAは、クロスディープな入力適応性、モダリティを意識し、低ランクのボトルネックを通じて効率的にパラメータ化できる軽量なメカニズムである。
- 参考スコア(独自算出の注目度): 55.74376789006731
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Supervised fine-tuning (SFT) on visual instruction data often improves perceptual capabilities in vision-language models (VLMs) while degrading reasoning performance, creating a persistent reasoning tax during post-training. We investigate whether this degradation is related to disrupted access to depth-wise representations, and find that even fixed cross-depth aggregation substantially restores reasoning, suggesting that preserved cross-depth access is an important missing factor in VLM fine-tuning. Building on this observation, we propose Input-Adaptive Depth Aggregation (IADA), a lightweight mechanism that makes cross-depth retrieval input-adaptive, modality-aware, and efficiently parameterized through a low-rank bottleneck. On Qwen3-VL-2B, IADA improves the average reasoning score by 9.5 points and the average perception score by $3.3$ points over LoRA-only fine-tuning with only 0.14M additional parameters, with the strongest gains appearing in parameter-efficient low-rank settings.
- Abstract(参考訳): 視覚的インストラクションデータに対する教師付き微調整(SFT)は、視覚言語モデル(VLM)の知覚能力を向上すると同時に、推論性能を低下させ、後トレーニング中に永続的な推論税を発生させる。
この劣化が深度表現の破壊的アクセスと関係しているかどうかを考察し、固定された深度集約でさえ推論を著しく復元することを発見し、保存された深度アクセスがVLMの微調整において重要な欠落要因であることを示唆した。
本研究は,入力-適応深度集約(IADA)を提案する。これは,クロスディープ入力を適応し,モダリティを意識し,低ランクのボトルネックを通じて効率的にパラメータ化する軽量な機構である。
Qwen3-VL-2Bでは、平均推論スコアを9.5ポイント改善し、平均認識スコアをわずか0.14MのパラメータでLoRAのみの微調整で3.3ドルとした。
関連論文リスト
- Smooth Operator: Smooth Verifiable Reward Activates Spatial Reasoning Ability of Vision-Language Model [18.526821056010384]
視覚言語モデル(VLM)は、3次元シーン理解のための正確な数値予測を実現する上で重要なボトルネックに直面している。
伝統的な強化学習アプローチは、主に相対的なランクに基づいており、しばしば深刻な報酬の分散と勾配不安定に悩まされる。
本稿では,Smooth Numerical Reward Activation (SNRA)演算子とAbsolute-Preserving GRPOフレームワークを紹介する。
論文 参考訳(メタデータ) (2026-01-12T16:26:42Z) - IMSE: Efficient U-Net-based Speech Enhancement using Inception Depthwise Convolution and Amplitude-Aware Linear Attention [2.3959703715401903]
本稿では,系統的に最適化された超軽量ネットワークIMSEを提案する。
1) MET モジュールを Amplitude-Aware Linear Attention (MALA) に、2) Deformable Embedding (DE) モジュールを Inception Depthwise Convolution (IDConv) に置き換える。
実験では、IMSEはパラメータ数を16.8%(0.513Mから0.427M)削減し、PESQ測定値(3.373)の最先端技術に匹敵する競争性能を達成する。
論文 参考訳(メタデータ) (2025-11-18T14:11:54Z) - Token-Level Inference-Time Alignment for Vision-Language Models [58.41370989069588]
VLM(Vision-Language Models)は、現代のマルチモーダルインテリジェンスの重要なバックボーンとなっている。
本稿では,基本VLMを凍結し,その分布を近似する報酬モデルをトレーニングする軽量フレームワークTITAを提案する。
推測中、暗黙の選好信号は報酬モデルと目標VLMの対数確率比として抽出され、密集した自己回帰フィードバックが得られる。
論文 参考訳(メタデータ) (2025-10-20T09:58:03Z) - Perception-Aware Policy Optimization for Multimodal Reasoning [79.56070395437898]
現在のマルチモーダル推論における誤りの主な原因は、視覚入力の知覚にある。
提案するPAPOは,モデルが推論を学習しながら知覚を学習することを奨励する,新しいポリシー勾配アルゴリズムである。
知覚誤りの30.5%が有意に減少し,PAPOによる知覚能力の向上が示唆された。
論文 参考訳(メタデータ) (2025-07-08T23:22:34Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。