論文の概要: FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers
- arxiv url: http://arxiv.org/abs/2607.14898v1
- Date: Thu, 16 Jul 2026 12:16:12 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-17 17:01:33.091571
- Title: FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers
- Title(参考訳): FlashDecoder: トランスフォーマーを備えたリアルタイム遅延-ピクセル間ストリーミングデコーダ
- Authors: Minguk Kang, Suha Kwak,
- Abstract要約: FlashDecoderは高速でメモリ効率の良い純粋なトランスフォーマービデオデコーダである。
最大1080pの解像度で、ラテントをピクセルフレームにデコードできる。
- 参考スコア(独自算出の注目度): 43.94286178727893
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Real-time video generation demands fast decoding as much as fast denoising, yet current latent video diffusion models rely on 3D convolutional decoders that are slow and memory-intensive at high resolutions or for long video. We introduce FlashDecoder, a fast, memory-efficient pure-Transformer video decoder that decodes latents to pixels frame by frame. At each step, the current frame attends only to a fixed-size window of past frames through a rolling KV cache. The fixed temporal window keeps decoding fast and memory bounded regardless of video length, enabling constant-latency streaming. Because frames are processed sequentially, temporal causality is enforced without explicit attention masks, enabling training at resolutions up to 1080p and matching the reconstruction quality of convolutional decoders. On the Wan2.1 and Wan2.2 latent spaces, FlashDecoder matches each convolutional decoder in reconstruction quality (e.g., 41.55dB vs. 41.49dB PSNR at 1080p) while decoding 3.6x-4.7x faster with up to 11x less memory on a single H100 GPU. With architecture-aware inference optimizations, the speedup widens to 12x.
- Abstract(参考訳): リアルタイムビデオ生成は高速デコードと高速デコードを必要とするが、現在の遅延ビデオ拡散モデルは3D畳み込みデコーダに依存している。
高速でメモリ効率のよいPure-TransformerビデオデコーダであるFlashDecoderを導入し、ラテントをフレーム単位でピクセルフレームにデコードする。
各ステップにおいて、現在のフレームは、ローリングKVキャッシュを介して過去のフレームの固定サイズのウィンドウにのみ対応する。
固定時間ウィンドウは、ビデオ長にかかわらずデコードとメモリのバウンドを保ち、一定のレイテンシのストリーミングを可能にする。
フレームは順次処理されるため、時間的因果性は明示的な注意マスクなしで実施され、最大1080pの解像度でのトレーニングを可能にし、畳み込みデコーダの復元品質と一致する。
Wan2.1 と Wan2.2 のラテント空間では、FlashDecoder は各畳み込みデコーダを再構成品質 (例: 41.55dB vs. 41.49dB PSNR at 1080p) で一致させ、一方のH100 GPUでは最大11倍のメモリで 3.6x-4.7x 高速に復号する。
アーキテクチャを意識した推論最適化により、スピードアップは12倍に拡大する。
関連論文リスト
- DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder [55.26098043655325]
DC-VideoGenは、事前訓練されたビデオ拡散モデルに適用することができる。
軽量な微調整を施した深部圧縮潜伏空間に適応することができる。
論文 参考訳(メタデータ) (2025-09-29T17:59:31Z) - SIEDD: Shared-Implicit Encoder with Discrete Decoders [36.705337163276255]
Inlicit Neural Representations (INR)は、ビデオごとの最適化機能を学ぶことによって、ビデオ圧縮に例外的な忠実度を提供する。
既存のINRエンコーディングの高速化の試みは、しばしば再建品質や重要な座標レベルの制御を犠牲にしている。
これらの妥協なしにINRエンコーディングを根本的に高速化する新しいアーキテクチャであるSIEDDを紹介する。
論文 参考訳(メタデータ) (2025-06-29T19:39:43Z) - QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design [54.38970077613728]
ビデオ監視、会議要約、教育講義分析、スポーツ放送といった現実の応用において、ロングビデオ理解が重要な機能として現れてきた。
我々は,リアルタイムダウンストリームアプリケーションをサポートするために,長時間ビデオ理解を大幅に高速化するシステムアルゴリズムの共同設計であるQuickVideoを提案する。
論文 参考訳(メタデータ) (2025-05-22T03:26:50Z) - Toward Lightweight and Fast Decoders for Diffusion Models in Image and Video Generation [0.0]
大きな変分オートエンコーダデコーダは、生成を遅くし、かなりのGPUメモリを消費することができる。
軽量なVision Transformer と Taming Transformer アーキテクチャを用いたカスタムトレーニングデコーダを提案する。
COCO 2017では、画像生成の全体的なスピードアップが最大15%、サブモジュールでのデコーディングが最大20倍、ビデオタスクのUCF-101がさらに向上している。
論文 参考訳(メタデータ) (2025-03-06T16:21:49Z) - Fast Encoding and Decoding for Implicit Video Representation [88.43612845776265]
本稿では,高速エンコーディングのためのトランスフォーマーベースのハイパーネットワークであるNeRV-Encと,効率的なビデオローディングのための並列デコーダであるNeRV-Decを紹介する。
NeRV-Encは勾配ベースの最適化をなくすことで$mathbf104times$の素晴らしいスピードアップを実現している。
NeRV-Decはビデオデコーディングを単純化し、ロード速度が$mathbf11times$で従来のコーデックよりも高速である。
論文 参考訳(メタデータ) (2024-09-28T18:21:52Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。