論文の概要: Motif-Video 2B: Technical Report
- arxiv url: http://arxiv.org/abs/2604.16503v1
- Date: Tue, 14 Apr 2026 15:09:39 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-04-21 21:52:52.034038
- Title: Motif-Video 2B: Technical Report
- Title(参考訳): Motif-Video 2B:テクニカルレポート
- Authors: Junghwan Lim, Wai Ting Cheung, Minsu Ha, Beomgyu Kim, Taewhan Kim, Haesol Lee, Dongpin Oh, Jeesoo Lee, Taehyun Kim, Minjae Kim, Sungmin Lee, Hyeyeon Cho, Dahye Choi, Jaeheui Her, Jaeyeon Huh, Hanbin Jung, Changjin Kang, Dongseok Kim, Jangwoong Kim, Youngrok Kim, Hyukjin Kweon, Hongjoo Lee, Jeongdoo Lee, Junhyeok Lee, Eunhwan Park, Yeongjae Park, Bokki Ryu, Dongjoo Weon,
- Abstract要約: 強力なビデオ生成モデルを訓練するには、通常、大量のデータセット、大きなパラメータ数、相当量の計算が必要である。
本研究では,1000万回未満のクリップと10万回のH200GPU時間という,より小さな予算で,強力なテキスト・ビデオ品質を実現することができるかどうかを問う。
- 参考スコア(独自算出の注目度): 10.684194955950353
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Training strong video generation models usually requires massive datasets, large parameter counts, and substantial compute. In this work, we ask whether strong text-to-video quality is possible at a much smaller budget: fewer than 10M clips and less than 100,000 H200 GPU hours. Our core claim is that part of the answer lies in how model capacity is organized, not only in how much of it is used. In video generation, prompt alignment, temporal consistency, and fine-detail recovery can interfere with one another when they are handled through the same pathway. Motif-Video 2B addresses this by separating these roles architecturally, rather than relying on scale alone. The model combines two key ideas. First, Shared Cross-Attention strengthens text control when video token sequences become long. Second, a three-part backbone separates early fusion, joint representation learning, and detail refinement. To make this design effective under a limited compute budget, we pair it with an efficient training recipe based on dynamic token routing and early-phase feature alignment to a frozen pretrained video encoder. Our analysis shows that later blocks develop clearer cross-frame attention structure than standard single-stream baselines. On VBench, Motif-Video~2B reaches 83.76\%, surpassing Wan2.1 14B while using 7$\times$ fewer parameters and substantially less training data. These results suggest that careful architectural specialization, combined with an efficiency-oriented training recipe, can narrow or exceed the quality gap typically associated with much larger video models.
- Abstract(参考訳): 強力なビデオ生成モデルを訓練するには、通常、大量のデータセット、大きなパラメータ数、相当量の計算が必要である。
本研究では,1000万回未満のクリップと10万回のH200GPU時間という,より小さな予算で,強力なテキスト・ビデオ品質を実現することができるかどうかを問う。
私たちの中核的な主張は、答えの一部はモデルキャパシティの組織化の仕方にある、ということです。
ビデオ生成では、プロンプトアライメント、時間的整合性、細部回復は、同じ経路で処理されるときに互いに干渉することがある。
Motif-Video 2Bは、スケールのみに頼るのではなく、これらの役割をアーキテクチャ的に分離することで、この問題に対処する。
モデルには2つの重要なアイデアが組み合わさっている。
まず,ビデオトークンシーケンスが長くなるとテキスト制御が強化される。
第二に、3つの部分からなるバックボーンは、初期の融合、共同表現学習、詳細改善を分離する。
この設計を限られた計算予算で効果的にするために、動的トークンルーティングと早期特徴アライメントに基づく効率的なトレーニングレシピと、凍結した事前学習ビデオエンコーダとのペアリングを行う。
分析により、後続のブロックは標準の単一ストリームベースラインよりも明瞭なクロスフレームアテンション構造を発達させることが示された。
VBench では、Motif-Video~2B は 83.76 % に達し、Wan2.1 14B を上回り、7$\times$ のパラメータを減らし、トレーニングデータを大幅に減らした。
これらの結果は、注意深いアーキテクチャの特殊化と効率志向のトレーニングレシピが組み合わさって、より大規模なビデオモデルに典型的な品質格差を狭めたり超えたりすることができることを示唆している。
関連論文リスト
- Generative Video Matting [57.186684844156595]
ビデオ・マッティングは、伝統的に高品質な地上データがないために制限されてきた。
既存のビデオ・マッティング・データセットのほとんどは、人間が注釈付けした不完全なアルファとフォアグラウンドのアノテーションのみを提供する。
本稿では,事前学習したビデオ拡散モデルから,よりリッチな事前処理を効果的に活用できる新しいビデオマッチング手法を提案する。
論文 参考訳(メタデータ) (2025-08-11T12:18:55Z) - A Simple Recipe for Contrastively Pre-training Video-First Encoders Beyond 16 Frames [57.758863967770594]
我々は,大規模な画像テキストモデルを浅部時間融合によりビデオに転送する共通パラダイムを構築した。
1)標準ビデオデータセットにおけるビデオ言語アライメントの低下による空間能力の低下と,(2)処理可能なフレーム数のボトルネックとなるメモリ消費の増大である。
論文 参考訳(メタデータ) (2023-12-12T16:10:19Z) - VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking [57.552798046137646]
Video masked autoencoder(ビデオマスクオートエンコーダ)は、ビデオ基礎モデルを構築するための、スケーラブルで汎用的な自己監督型プレトレーナーである。
我々は10億のパラメータを持つビデオViTモデルのトレーニングに成功した。
論文 参考訳(メタデータ) (2023-03-29T14:28:41Z) - VA-RED$^2$: Video Adaptive Redundancy Reduction [64.75692128294175]
我々は,入力依存の冗長性低減フレームワークva-red$2$を提案する。
ネットワークの重み付けと協調して適応ポリシーを共有重み付け機構を用いて微分可能な方法で学習する。
私たちのフレームワークは、最先端の方法と比較して、計算(FLOP)の20% - 40%$削減を達成します。
論文 参考訳(メタデータ) (2021-02-15T22:57:52Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。