論文の概要: Aggregating Visual Information with Optimal Transport for VideoLM Token Compression
- arxiv url: http://arxiv.org/abs/2608.20473v2
- Date: Tue, 25 Aug 2026 09:52:24 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-26 14:09:33.82451
- Title: Aggregating Visual Information with Optimal Transport for VideoLM Token Compression
- Title(参考訳): 動画像圧縮のための最適移動による視覚情報の集約
- Authors: Wenti Yin, Xiaotian Han, Junyuan Shang, Yuchen Ding, Shuohuan Wang, Dianhai Yu, Changxin Gao, Nong Sang,
- Abstract要約: ビデオ言語モデルは、ビデオを表現の冗長性に富んだ濃密な視覚的なシーケンスとして処理する。
これらのシーケンスは、言語モデルデコーディングにおける視覚的負担を軽減するために不可欠である。
AVIOT(Aggregating Visual Information with Optimal Transport)を導入し,フレーム観測の高密度な経験的指標としてビデオトークン圧縮を論じる。
- 参考スコア(独自算出の注目度): 55.867497983465476
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Video language models process videos as dense visual-token sequences with substantial representational redundancy. Compressing these sequences is therefore essential for reducing the visual-token burden on language-model decoding. The central challenge is to preserve visual information dispersed across frames under such compression. To this end, we introduce Aggregating Visual Information with Optimal Transport (AVIOT), which casts video token compression as transporting a dense empirical measure of frame observations onto a compact target measure. The resulting source-to-target coupling induces a distribution over source observations for each target support, directly specifying how the compressed video representation is constructed. We further adapt this construction along task and spatial axes. Question conditioning modulates the transport cost between source frames and target supports, while influencing how many supports are allocated to each temporal segment, thereby directing representation capacity toward question-relevant content. At multiple spatial granularities, AVIOT computes region-specific temporal transport plans and adaptively fuses the representations they yield, allowing different regions within the same compact representation to draw from different moments. Evaluations across varying compression ratios show that AVIOT matches or outperforms the uncompressed baseline on multiple video-understanding benchmarks while retaining strong performance at higher compression ratios.
- Abstract(参考訳): ビデオ言語モデルは、ビデオを表現の冗長性に富んだ濃密な視覚的なシーケンスとして処理する。
したがって、これらのシーケンスを圧縮することは、言語モデルデコーディングにおける視覚的負担を軽減するために不可欠である。
中心的な課題は、そのような圧縮の下でフレーム全体に分散した視覚情報を保存することである。
この目的のために,AVIOT(Aggregating Visual Information with Optimal Transport)を提案する。
得られたソース・トゥ・ターゲット結合は、圧縮された映像表現がどのように構成されているかを直接指定して、各目標支援のソース観測上の分布を誘導する。
我々はこの構造をさらにタスクと空間軸に沿って順応する。
質問条件付けは、ソースフレームとターゲットサポート間の転送コストを調整し、各時間セグメントにサポート数に影響を与えることにより、質問関連コンテンツへの表現能力を誘導する。
複数の空間的粒度において、AVIOTは地域固有の時間的輸送計画(英語版)を計算し、それらが生成する表現を適応的に融合させ、同じコンパクトな表現内の異なる領域が異なるモーメントから引き出すことを可能にする。
様々な圧縮比による評価では、AVIOTはより高い圧縮比で高い性能を維持しながら、複数のビデオアンダースタンドドベンチマークにおいて圧縮されていないベースラインと一致または性能を向上する。
関連論文リスト
- Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding [49.540751364954666]
Visual Token Codecは、グローバルトークンとパッチトークンを専用のコーディングパスに分離する。
VTCは, 分類, セグメンテーション, 検出タスクにおいて, 代表的VT特徴量ベースラインを一貫して上回っていることを示す。
論文 参考訳(メタデータ) (2026-08-09T17:32:04Z) - OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models [39.853488828881986]
OTT-Vidは、時間的トークン圧縮のためのトランスポートから派生したアロケーションフレームワークである。
OTT-VidはVQAの95.8%、VTGのパフォーマンスの73.9%を維持し、トークンの10%しか保持していない。
論文 参考訳(メタデータ) (2026-05-12T08:58:49Z) - FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding [55.700832127331324]
FLoCは、施設位置関数に基づく効率的なビジュアルトークン圧縮フレームワークである。
本手法は,トークンのコンパクトな部分集合を迅速に選択することにより,顕著な効率向上を実現する。
私たちのアプローチは、トレーニング不要、モデル非依存、クエリ非依存で、汎用的なソリューションを提供しています。
論文 参考訳(メタデータ) (2025-10-31T17:29:39Z) - State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video Understanding [50.866929044215965]
本稿では,映像理解のためのステートスペース・プロンプティング(SSP)手法を提案する。
SSPはフレーム内のプロンプトを組み合わせて、ビデオ内の重要な時間情報を集約し、伝達する。
我々のSSPは、既存のSOTA法を平均2.76%上回っている。
論文 参考訳(メタデータ) (2025-10-14T05:30:36Z) - Embedding Compression Distortion in Video Coding for Machines [67.97469042910855]
現在、ビデオ伝送は人間の視覚システム(HVS)だけでなく、分析のための機械認識にも役立っている。
本稿では,機械知覚関連歪み表現を抽出し,下流モデルに埋め込む圧縮歪埋め込み(CDRE)フレームワークを提案する。
我々のフレームワークは,実行時間,パラメータ数といったオーバーヘッドを最小限に抑えて,既存のコーデックのレートタスク性能を効果的に向上させることができる。
論文 参考訳(メタデータ) (2025-03-27T13:01:53Z) - PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models [64.9366388601049]
ビジュアルトークン圧縮は、視覚入力の相当なトークン長を減らすために利用される。
我々は,プログレッシブ・ビジュアル・トークン圧縮と呼ばれる統一的なトークン圧縮戦略を導入する。
本モデルは,様々なビデオ理解ベンチマークにおいて,最先端のパフォーマンスを実現する。
論文 参考訳(メタデータ) (2024-12-12T18:59:40Z) - Neural-based Video Compression on Solar Dynamics Observatory Images [8.73521037463594]
NASAのソーラー・ダイナミクス・オブザーバトリー(SDO)ミッションは、太陽の日常活動を監視するために膨大なデータを収集する。
データ圧縮は、限られたテレメトリレートによって引き起こされる課題に対処する上で重要な役割を果たす。
本稿では,SDOの画像データ収集における圧縮率の高いニューラルビデオ圧縮手法を提案する。
論文 参考訳(メタデータ) (2024-07-12T21:24:25Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。