論文の概要: VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders
- arxiv url: http://arxiv.org/abs/2607.14088v1
- Date: Wed, 15 Jul 2026 17:59:23 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-16 16:39:12.857131
- Title: VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders
- Title(参考訳): VideoRAE: 表現オートエンコーダによる生成モデリングのためのビデオファウンデーションモデル
- Authors: Zhihao Xie, Junfeng Wu, Xinting Hu, Junchao Huang, Li Jiang,
- Abstract要約: VJEPAやVideoMAEv2のようなビデオファンデーションモデル(VF-Ms)は、強力なビデオ理解能力を示している。
VideoRAEは、凍結したビデオファウンデーションエンコーダの階層的特徴を活用して、軽量な1Dセルフアテンションで圧縮する表現オートエンコーダである。
実験の結果, VideoRAEは連続型と離散型の両方で強い再構成が達成された。
- 参考スコア(独自算出の注目度): 14.313320718526414
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Video generative models commonly rely on latent spaces learned by 3D Variational Autoencoders (3D-VAEs). However, conventional 3D-VAEs are mainly optimized for pixel-level reconstruction, which can limit the semantic and spatio-temporal structure captured by their latents. Meanwhile, Video Foundation Models (VFMs) such as V-JEPA 2 and VideoMAEv2 show strong video understanding capabilities, yet whether their frozen representations can be transformed into compact, reconstruction-capable, and generation-friendly video latents remains largely unexplored. We answer this question with VideoRAE, a representation autoencoder that leverages multi-scale hierarchical features from a frozen video foundation encoder and compresses them with a lightweight 1D self-attention projector. VideoRAE supports both continuous latents for Diffusion Transformers and discrete tokens for autoregressive models via multi-codebook high-dimensional quantization. During decoding, a local-and-global representation alignment objective with the frozen VFM teacher improves semantic preservation and enables training without KL regularization. Experiments show that VideoRAE achieves strong reconstruction in both continuous and discrete regimes. On UCF-101, it obtains state-of-the-art class-to-video gFVDs of 40 and 93 with AR and DiT generators, respectively, while converging approximately 5x faster than competing autoencoder baselines. In a controlled 2B-scale text-to-video study, replacing LTX-VAE with VideoRAE leads to faster convergence under comparable settings. These results validate frozen VFM representations as versatile and generation-friendly video latents. The model and code will be released on https://zhxie0117.github.io/VideoRAE.
- Abstract(参考訳): ビデオ生成モデルは一般に3D変分オートエンコーダ(3D-VAEs)によって学習される潜在空間に依存している。
しかし,従来の3D-VAEは,主に画素レベルの再構成に最適化されている。
一方、V-JEPA 2 や VideoMAEv2 のようなビデオファウンデーションモデル(VFM)は、強力なビデオ理解能力を示しているが、その凍結表現がコンパクトで再構成可能で、世代フレンドリーなビデオラテントに変換できるかどうかはまだ明らかになっていない。
我々は,凍結したビデオファウンデーションエンコーダのマルチスケール階層的特徴を活用し,軽量な1次元自己保持プロジェクタで圧縮する表現オートエンコーダであるVideoRAEで,この問題に答える。
VideoRAEは拡散変換器の連続潜時とマルチコードブックの高次元量子化による自己回帰モデルの離散トークンの両方をサポートする。
復号中、凍結したVFM教師による局所的・言語的表現アライメントの目的は、意味保存を改善し、KL正規化を伴わない訓練を可能にする。
実験の結果, VideoRAEは連続型と離散型の両方で強い再構成が達成された。
UCF-101では、ARとDiTジェネレータをそれぞれ40と93の最先端のクラス対ビデオのgFVDを取得し、競合するオートエンコーダのベースラインよりも約5倍高速に収束する。
LTX-VAEをVideoRAEに置き換える2Bスケールのテキスト・ツー・ビデオの研究は、同等の設定下でより高速な収束をもたらす。
これらの結果は,VFMの凍結表現を多目的かつ世代フレンドリーなビデオラテントとして検証した。
モデルとコードはhttps://zhxie0117.github.io/VideoRAEでリリースされる。
関連論文リスト
- Towards Holistic Modeling for Video Frame Interpolation with Auto-regressive Diffusion Transformers [95.68243351895107]
我々はtextbfVideo textbfFrame textbfInterpolation (LDF-VFI) のための textbfLocal textbfDiffusion textbfForcing for textbfVideo textbfFrame textbfInterpolation (LDF-VFI) という包括的でビデオ中心のパラダイムを提案する。
我々のフレームワークは、ビデオシーケンス全体をモデル化し、長距離時間的コヒーレンスを確保する自動回帰拡散変換器上に構築されている。
LDF-VFIは、挑戦的なロングシーケンスベンチマークで最先端のパフォーマンスを達成する。
論文 参考訳(メタデータ) (2026-01-21T12:58:52Z) - MotionAura: Generating High-Quality and Motion Consistent Videos using Discrete Diffusion [3.7270979204213446]
ビデオ処理の課題に対処するための4つの重要なコントリビューションを提示する。
まず,3次元逆ベクトル量子化バリエンコエンコオートコーダを紹介する。
次に,テキスト・ビデオ生成フレームワークであるMotionAuraを紹介する。
第3に,スペクトル変換器を用いたデノナイジングネットワークを提案する。
第4に,Sketch Guided Videopaintingのダウンストリームタスクを導入する。
論文 参考訳(メタデータ) (2024-10-10T07:07:56Z) - When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding [118.72266141321647]
CMVC(Cross-Modality Video Coding)は、ビデオ符号化における多モード表現とビデオ生成モデルを探索する先駆的な手法である。
復号化の際には、以前に符号化されたコンポーネントとビデオ生成モデルを利用して複数の復号モードを生成する。
TT2Vは効果的な意味再構成を実現し,IT2Vは競争力のある知覚整合性を示した。
論文 参考訳(メタデータ) (2024-08-15T11:36:18Z) - Hybrid Video Diffusion Models with 2D Triplane and 3D Wavelet Representation [35.52770785430601]
複雑な依存関係をより効率的にキャプチャできるHVtemporalDMというハイブリッドビデオオートエンコーダを提案する。
HVDMは、ビデオの歪んだ表現を抽出するハイブリッドビデオオートエンコーダによって訓練される。
当社のハイブリッドオートエンコーダは、生成されたビデオに詳細な構造と詳細を付加した、より包括的なビデオラテントを提供します。
論文 参考訳(メタデータ) (2024-02-21T11:46:16Z) - RAVEN: Rethinking Adversarial Video Generation with Efficient Tri-plane Networks [93.18404922542702]
本稿では,長期的空間的および時間的依存関係に対処する新しいビデオ生成モデルを提案する。
提案手法は,3次元認識型生成フレームワークにインスパイアされた,明示的で単純化された3次元平面のハイブリッド表現を取り入れたものである。
我々のモデルは高精細度ビデオクリップを解像度256時間256$ピクセルで合成し、フレームレート30fpsで5ドル以上まで持続する。
論文 参考訳(メタデータ) (2024-01-11T16:48:44Z) - Video Probabilistic Diffusion Models in Projected Latent Space [75.4253202574722]
我々は、PVDM(Latent Video diffusion model)と呼ばれる新しいビデオ生成モデルを提案する。
PVDMは低次元の潜伏空間で映像配信を学習し、限られた資源で高解像度映像を効率的に訓練することができる。
論文 参考訳(メタデータ) (2023-02-15T14:22:34Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。