論文の概要: FlashGaze: Training-Free Multi-Scale Patch Pruning For Efficient Video Understanding
- arxiv url: http://arxiv.org/abs/2610.04225v1
- Date: Sat, 03 Oct 2026 02:33:01 GMT
- ステータス: 情報取得中
- システム内更新日: 2026-10-06 21:23:17.430214
- Title: FlashGaze: Training-Free Multi-Scale Patch Pruning For Efficient Video Understanding
- Title(参考訳): FlashGaze:効率的なビデオ理解のための訓練不要なマルチスケールパッチ
- Abstract要約: FlashGazeは、補助的なネットワークを導入することなくエンコーディング前に時間的冗長性を減少させる、トレーニング不要の方法である。
複数のベンチマークにまたがる2つのMLLMバックボーンの実験は、ほぼ精度を保ちながら、かなりの効率向上を示した。
これらの効率向上により、同じGPUハードウェア上で、より多くのフレームと高解像度の動画を処理できるようになる。
- 参考スコア(独自算出の注目度): 4.42290737383658
- License:
- Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong performance in video understanding, yet efficiently processing long, high-resolution videos remains challenging. Such videos often contain substantial spatiotemporal redundancy, and processing redundant visual tokens can incur avoidable computational overhead. Many existing methods prune visual tokens during or after vision transformer (ViT) encoding, leaving much of the encoding cost unaddressed. Some approaches prune patches before encoding but rely on learned auxiliary networks for patch selection, incurring additional training and inference overhead. To address these limitations, we propose FlashGaze, a training-free method that reduces spatiotemporal redundancy before ViT encoding without introducing auxiliary networks. FlashGaze uses pixel-space differences as a proxy for information loss and employs Quadtree Dynamic Programming to jointly optimize patch dropping, merging, and keeping under a fixed budget. Experiments on two MLLM backbones across multiple benchmarks demonstrate substantial efficiency gains while largely preserving accuracy. On Qwen3-VL-8B, FlashGaze retains 98% of the full-input baseline accuracy on LongVideoBench while achieving up to 5.4x and 17x speedups in ViT encoding and MLLM prefill, respectively, and reducing peak GPU memory usage by a factor of 1.8. These efficiency gains enable the model to process videos with more frames and higher resolutions on the same GPU hardware, unlocking video understanding at scales previously out of reach.
- Abstract(参考訳): MLLM(Multimodal Large Language Models)は、ビデオ理解において高い性能を示すが、長大な高解像度の動画を効率的に処理することは依然として困難である。
このようなビデオは時空間のかなりの冗長性を含み、冗長な視覚トークンの処理は回避可能な計算オーバーヘッドを発生させる可能性がある。
多くの既存の方法で視覚変換器 (ViT) の符号化中に視覚トークンをプルークし、符号化コストの大半を未調整のまま残している。
いくつかのアプローチでは、エンコーディング前にパッチをプルーンするが、パッチ選択のために学習した補助ネットワークに依存し、追加のトレーニングと推論オーバーヘッドが発生する。
これらの制約に対処するため、補助的なネットワークを導入することなく、ViTエンコーディング前の時空間冗長性を低減できる訓練不要のFlashGazeを提案する。
FlashGazeは、情報損失のプロキシとしてピクセル空間の違いを使用し、パッチドロップ、マージ、固定予算の維持を共同で最適化するためにQuadtree Dynamic Programmingを使用している。
複数のベンチマークにまたがる2つのMLLMバックボーンの実験は、ほぼ精度を保ちながら、かなりの効率向上を示した。
Qwen3-VL-8Bでは、FlashGazeはLongVideoBenchのフルインプットベースライン精度の98%を保持し、ViTエンコーディングとMLLMプリフィルの最大5.4倍と17倍のスピードアップを実現し、ピークGPUメモリ使用量を1.8倍に削減している。
これらの効率向上により、同じGPUハードウェア上で、より多くのフレームと高解像度の動画を処理できるようになる。
関連論文リスト
- LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs [90.77662862634509]
LiteFrameは、ビデオ大言語モデルのための強力な、しかし非常に効率的なバックボーンである。
LiteFrameはエンドツーエンドのレイテンシを35%削減し、8$times$より多くのフレームを処理する。
計算予算の固定化により,より長めの映像理解を解き明かす可能性を示した。
論文 参考訳(メタデータ) (2026-05-17T05:02:52Z) - STORM: Token-Efficient Long Video Understanding for Multimodal LLMs [116.4479155699528]
STORMは、イメージエンコーダとビデオLLMの間に専用のテンポラリエンコーダを組み込んだ、新しいアーキテクチャである。
我々は,STORMが様々な長いビデオ理解ベンチマークにおいて最先端の結果を達成することを示す。
論文 参考訳(メタデータ) (2025-03-06T06:17:38Z) - SparseTem: Boosting the Efficiency of CNN-Based Video Encoders by Exploiting Temporal Continuity [19.900719882624028]
本稿では,メモリオーバーヘッドを削減するためのメモリ効率スケジューリング手法と,精度の劣化を最小限に抑えるためのオンライン調整機構を提案する。
SparseTemは効率の良いDetでは1.79x、CRNNでは4.72xの高速化を実現している。
論文 参考訳(メタデータ) (2024-10-28T07:13:25Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。