論文の概要: MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs
- arxiv url: http://arxiv.org/abs/2610.01434v1
- Date: Thu, 01 Oct 2026 10:31:21 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-03 01:19:24.056693
- Title: MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs
- Title(参考訳): MWOP: 効率的なMLLMのためのモダリティ対応幅ワイズ操作
- Abstract要約: MWOP(Modality-Aware Width-wise Operation Pruning)を提案する。
MWOP prunes visual-to-visual (V2V), text-to-visual (T2V), text-to-text (T2T) attention paths in each layer, and selects FFN channel for visual and textual inputs。
トークン列は、注目度とFFN計算を低減しつつ保存し、トークン圧縮を補完し、シーケンス長とトークン単位の計算を同時に削減する。
- 参考スコア(独自算出の注目度): 24.41434670466578
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Multimodal large language models (MLLMs) incur substantial inference costs when processing long visual-textual sequences. While existing operation compression methods exploit modality-level redundancy, they largely treat computation within attention heads and shared feed-forward network (FFN) channels as unified units, leaving finer-grained redundancy underexplored. We find that redundancy varies both across modality-interaction paths within the same attention head and across visual and textual executions of the same FFN channel. Based on these findings, we propose Modality-aware Width-wise Operation Pruning (MWOP), which independently prunes visual-to-visual (V2V), text-to-visual (T2V), and text-to-text (T2T) attention paths within each layer, and separately selects FFN channels for visual and textual inputs. A first-order Taylor criterion guides the pruning process, with FFN importance re-evaluated after attention pruning and LoRA-based recovery training. To translate the resulting fine-grained sparsity into practical acceleration, we further develop path-sparse Triton attention kernels and compact visual-side FFN execution. MWOP preserves the token sequence while reducing attention and FFN computation, making it complementary to token compression and enabling simultaneous reduction of sequence length and per-token computation. On LLaVA-OneVision-7B, MWOP alone achieves a $1.6\times$ prefill speedup with 99.7\% average performance retention across 12 benchmarks. Combined with two representative token compression methods, it further increases their prefill speedups from $2.0\times$ and $1.9\times$ to $2.9\times$ and $2.7\times$, respectively. Results on Qwen2.5-VL-7B further demonstrate its applicability across architectures. The code is available at https://github.com/EIT-NLP/MWOP.
- Abstract(参考訳): MLLM(Multimodal large language model)は、長い視覚・テクスチャシーケンスを処理する際に、かなりの推論コストを発生させる。
既存の運用圧縮手法では、モダリティレベルの冗長性を利用するが、注意頭と共有フィードフォワードネットワーク(FFN)チャネル内の計算を統一単位として扱い、よりきめ細かい冗長性の研究が過小評価されている。
冗長性は同じアテンションヘッド内のモダリティ-相互作用経路と、同じFFNチャネルの視覚的およびテキスト的実行の両方で変化する。
これらの知見に基づいて,視覚的・視覚的(V2V),テキスト的(T2V),テキスト的(T2T)の注意経路を独立に抽出するMWOP(Modality-Aware Width-wise Operation Pruning)を提案し,視覚的・テキスト的入力のためのFFNチャネルを別々に選択する。
1階目のTaylor criterion はプルーニングプロセスのガイドであり、FFN の重要性は注意プルーニングと LoRA ベースのリカバリトレーニングによって再評価される。
得られた微細な空間を実用的な加速度に変換するため、パススパースなトリトンアテンションカーネルと、コンパクトなビジュアルサイドFFN実行を開発する。
MWOPは、注目度とFFN計算を低減しつつトークンシーケンスを保存し、トークン圧縮を補完し、シーケンス長とトーケン毎の計算の同時削減を可能にする。
LLaVA-OneVision-7Bでは、MWOPだけで1.6\times$プリフィルのスピードアップを達成し、12ベンチマークの平均パフォーマンス保持率は99.7\%である。
2つの代表的なトークン圧縮法と組み合わせて、プリフィルのスピードアップを$2.0\times$と$1.9\times$から$2.9\times$と$2.7\times$に拡大する。
Qwen2.5-VL-7Bの結果はさらにアーキテクチャ間の適用性を示した。
コードはhttps://github.com/EIT-NLP/MWOPで公開されている。
関連論文リスト
- SPIDER: Multi-Layer Semantic Token Pruning and Adaptive Sub-Layer Skipping in Multimodal Large Language Models [108.54713532551811]
マルチレイヤのアンダーラインtextbfSemanticビジュアルトークン underlinetextbfPrununderlinetextbfIng と aunderlinetextbfDaptive sub-layunderlinetextbfER スキップ機構を統合したトレーニングフリーフレームワーク textbfSPIDER を提案する。
論文 参考訳(メタデータ) (2026-09-28T11:54:41Z) - AdaVSkip: Adaptive Visual Token Skipping Across Layers For Efficient MLLMs Inference [14.50173720268662]
マルチモーダル・大規模言語モデル(MLLM)は、すべてのトランスフォーマー層にまたがる多数の視覚トークンを処理するために、かなりの計算を必要とする。
本稿では,AdaVSkipを提案する。AdaVSkipは各層に2つの軽量ルータを装備し,視覚トークンが通過するか否かを独立に判断する。
AdaVSkipは、計算量を大幅に減らしながら、強いタスク性能を維持している。
論文 参考訳(メタデータ) (2026-09-14T07:07:51Z) - OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond [50.440302567029654]
マルチモーダルインテリジェンスにより、Key-Valueキャッシュは効率的なデプロイメントのための主要なメモリボトルネックとなった。
本研究では、チャネルごとの量子化パラダイムの本質的な限界を再考する。
X-LLMのための高精度かつ軽量なKVキャッシュ圧縮フレームワークOScaRを提案する。
論文 参考訳(メタデータ) (2026-05-19T10:53:03Z) - PIO-FVLM: Rethinking Training-Free Visual Token Reduction for VLM Acceleration from an Inference-Objective Perspective [59.24570811503256]
本稿では,視覚モデル(VLM)における冗長な視覚トークンを減らし,推論を高速化するPIO-FVLMを提案する。
提案されているPIO-FVLMは、トレーニングフリーで、FlashAttentionと互換性があり、実用的なアプリケーションやデプロイメントに親しみやすい。
LLaVA-Next-7Bでは、PIO-FVLMは視覚トークンの11.1%しか保持していないが、オリジナルのパフォーマンスの97.2%を維持している。
論文 参考訳(メタデータ) (2026-02-04T15:33:10Z) - HoliTom: Holistic Token Merging for Fast Video Large Language Models [32.620504076794795]
ビデオ言語モデル(ビデオLLM)は、ビデオ理解において優れるが、冗長なビデオトークンによる計算不効率に直面する。
HoliTomは、新しいトレーニング不要な全体的トークンフレームワークである。
また,内部LLMトークンの類似性に基づくマージ手法を導入する。
論文 参考訳(メタデータ) (2025-05-27T15:28:45Z) - RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs [38.34856927170692]
MLLM(Multimodal Large Language Model)の学習用フレームワークを提案する。
Probe-Activated Dynamic FFNとHollow Attentionで構成されており、ビジュアルトークンの計算の調整可能な削減を可能にする。
実験では、デコーダのみのMLLMに特有の、実質的で、構造化され、クラスタ化された冗長性を示す。
論文 参考訳(メタデータ) (2025-01-31T11:09:16Z) - An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models [65.37846460916042]
視覚的トークンに対する注意計算は,LVLMの深い層において極めて非効率であることがわかった。
本稿では,計算効率の最適化を目的とした多用途プラグアンドプレイ方式であるFastVを紹介する。
論文 参考訳(メタデータ) (2024-03-11T14:35:32Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。