論文の概要: CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation
- arxiv url: http://arxiv.org/abs/2607.03803v1
- Date: Sat, 04 Jul 2026 10:24:35 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-07 22:26:29.726176
- Title: CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation
- Title(参考訳): CineMobile:シネマカメラモーション生成のためのデバイス上の画像とビデオの拡散
- Authors: Xuyao Huang, Zelai Deng, Xu Wang, Xizhong Xiao, Zhijie Deng,
- Abstract要約: CineMobileは49フレームの480pビデオを生成し、NVIDIA H200 GPUでは0.6秒、MediaTek Dimensity 8400 Ultimate 5Gプラットフォームでは20秒を遅延する。
CineMobileは、ビジュアル品質を同等に保ちながら、世代別40倍のスピードアップを実現している。
- 参考スコア(独自算出の注目度): 17.654969014220033
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: The growing demand for image-to-video creation on mobile devices has increasingly focused on cinematic motion effects like bullet time, dolly zoom, slow motion, etc. While Diffusion Transformers (DiTs) exhibit strong performance in video generation, their large parameter sizes and multi-step iterative denoising processes lead to substantial computational overhead, making efficient generation on mobile devices challenging. We propose CineMobile to bridge the gap. In particular, CineMobile adopts a three-fold optimization strategy: (1) leveraging a distillation-guided pruning approach to derive a compact yet efficient model that retains the essential video generation capabilities required for cinematic effects; (2) optimizing the compressed model into a 4-step generator via a combination of diffusion distillation and reinforcement learning; (3) employing a hybrid post-training quantization strategy to compress the model footprint to under 1 GB. Experimental results show that compared to the teacher model with the Wan 2.1 architecture, CineMobile achieves a 40x speedup in generation while maintaining comparable visual quality. Specifically, CineMobile generates 49-frame 480p videos with a per-step denoising latency of 0.6s on an NVIDIA H200 GPU and 20s on the MediaTek Dimensity 8400 Ultimate 5G platform, with a peak memory usage of 1.8 GB, demonstrating its practical applicability for mobile-based image-to-video creation.
- Abstract(参考訳): モバイルデバイスにおける画像とビデオの制作に対する需要は、弾道時間、ドリーズーム、スローモーションなど、映画的な効果にますます焦点を絞っている。
Diffusion Transformers (DiTs) はビデオ生成において高い性能を示すが、その大きなパラメータサイズと多段階反復デノゲーションプロセスは計算オーバーヘッドを大幅に増加させ、モバイルデバイス上で効率的な生成を困難にしている。
我々はそのギャップを埋めるためにCineMobileを提案する。
特に、CineMobileは、(1)蒸留誘導プルーニング手法を利用して、撮影効果に必要なビデオ生成能力を維持するコンパクトで効率的なモデルを導出する、(2)蒸留蒸留と強化学習を組み合わせた4段階のジェネレータに圧縮モデルを最適化する、(3)ハイブリッドポストトレーニング量子化戦略を用いてモデルフットプリントを1GB未満に圧縮する、という3段階の最適化戦略を採用している。
実験の結果,教師モデルと Wan 2.1 アーキテクチャを比較すると,CineMobile は視覚的品質を同等に保ちながら,40倍の高速化を実現していることがわかった。
具体的には、CineMobileは49フレームの480pビデオを生成し、NVIDIA H200 GPUで0.6秒、MediaTek Dimensity 8400 Ultimate 5Gプラットフォームで20秒、ピークメモリが1.8GBで、モバイルベースのイメージ・ツー・ビデオ作成に実用的な適用性を示している。
関連論文リスト
- MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices [42.00270347221752]
モバイル端末上でのリアルタイム画像・ビデオ生成のための270M軽量拡散モデルであるMobileI2Vを提案する。
I2Vサンプリング工程を20回以上から2回まで圧縮する時間段階蒸留方式を設計した。
MobileI2Vは、モバイル端末で720pの高速動画生成を可能にする。
論文 参考訳(メタデータ) (2025-11-26T15:09:02Z) - Taming Diffusion Transformer for Efficient Mobile Video Generation in Seconds [91.56929670753226]
Diffusion Transformer (DiT) はビデオ生成タスクにおいて高いパフォーマンスを示しているが、その高い計算コストは、スマートフォンのようなリソース制約のあるデバイスでは実用的ではない。
本稿では,ビデオ生成の大幅な高速化と,モバイルプラットフォームへの実用的な展開を実現するための新しい最適化手法を提案する。
論文 参考訳(メタデータ) (2025-07-17T17:59:10Z) - SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device [61.42406720183769]
本稿では,大規模ビデオ拡散モデルのパワーをエッジユーザーにもたらすための包括的加速フレームワークを提案する。
我々のモデルは0.6Bのパラメータしか持たないため、iPhone 16 PMで5秒以内に5秒のビデオを生成することができる。
論文 参考訳(メタデータ) (2024-12-13T18:59:56Z) - Adaptive Caching for Faster Video Generation with Diffusion Transformers [52.73348147077075]
拡散変換器(DiT)はより大きなモデルと重い注意機構に依存しており、推論速度が遅くなる。
本稿では,Adaptive Caching(AdaCache)と呼ばれる,ビデオDiTの高速化のためのトレーニング不要手法を提案する。
また,AdaCache内で動画情報を利用するMoReg方式を導入し,動作内容に基づいて計算割り当てを制御する。
論文 参考訳(メタデータ) (2024-11-04T18:59:44Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。