論文の概要: Wan-Animate-2: Pushing the Application Boundaries of Character Animation
- arxiv url: http://arxiv.org/abs/2608.06009v1
- Date: Thu, 06 Aug 2026 13:13:10 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-07 15:25:20.911641
- Title: Wan-Animate-2: Pushing the Application Boundaries of Character Animation
- Title(参考訳): Wan-Animate-2: 文字アニメーションのアプリケーション境界を押す
- Abstract要約: Wan-Animate-2は、Diffusion Transformerで動画を直接消費するエンドツーエンドのキャラクターアニメーションフレームワークである。
本アーキテクチャは,中間運動抽出器を完全に取り除き,動きの忠実度とアイデンティティの保存性が向上する。
We present Wan-Animate-2-Lite, a efficientvariant that inference latency to real-time thresholds。
- 参考スコア(独自算出の注目度): 25.315356486955565
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Character image animation remains a foundational yet challenging task in computer vision. Existing approaches can be broadly categorized into three paradigms: methods based on explicit motion representations suffer from extraction errors and identity drift; methods based on implicit motion features lose fine-grained dynamics through compression; and in-context learning approaches avoid intermediate representations but incur prohibitive computational costs. Furthermore, all current systems are designed for offline synthesis, unable to meet the real-time requirements of interactive applications such as digital avatars and live-streaming hosts. To address these limitations, we present Wan-Animate-2, an end-to-end character animation framework that directly consumes the driving video within a redesigned Diffusion Transformer. Our architecture achieves superior motion fidelity and identity preservation by eliminating intermediate motion extractors entirely. We further introduce text driven viewpoint control that decouples the output camera perspective from the driving video--a capability rarely supported by prior character animation methods that rely on explicit motion representations. Beyond generation quality, we present Wan-Animate-2-Lite, an efficient variant that reduces inference latency to real-time thresholds through a three-stage training paradigm: teacher forcing pretraining with error buffer mechanism, and Self-Forcing distillation with chunk-wise backpropagation. This enables streaming character animation for interactive applications, opening new deployment scenarios that were previously infeasible. Qualitative evaluations and user studies demonstrate that Wan-Animate-2 achieves high-fidelity animation results across diverse characters and motion patterns. To foster further research and community development, we will release the Wan-Animate-2-Base model weights to the public.
- Abstract(参考訳): キャラクタイメージアニメーションは、コンピュータビジョンにおける基礎的かつ挑戦的な課題である。
既存のアプローチは3つのパラダイムに分類できる: 明示的な動作表現に基づく手法は抽出エラーやアイデンティティドリフトに苦しむ; 暗黙的な動作特徴に基づく手法は圧縮によってきめ細かなダイナミクスを失う; 文脈内学習アプローチは中間表現を避け、禁忌的な計算コストを抑える。
さらに、現在のシステムはすべてオフライン合成用に設計されており、デジタルアバターやライブストリーミングホストのようなインタラクティブなアプリケーションのリアルタイム要求を満たすことができない。
これらの制限に対処するため、我々は、Diffusion Transformerで直接動画を消費するエンドツーエンドのキャラクターアニメーションフレームワークであるWan-Animate-2を提案する。
本アーキテクチャは,中間運動抽出器を完全に取り除き,動きの忠実度とアイデンティティの保存性が向上する。
さらに、出力カメラの視点を駆動映像から切り離すテキスト駆動視点制御を導入する。
Wan-Animate-2-Liteは3段階のトレーニングパラダイムによって推論遅延をリアルタイムしきい値に低減し,エラーバッファ機構による事前学習,チャンクワイズバックプロパゲーションによる自己強制蒸留を行う。
これによりインタラクティブなアプリケーションのためのストリーミングキャラクタアニメーションが可能になり、以前は実現不可能だった新たなデプロイメントシナリオがオープンされる。
定性的評価とユーザスタディにより,Wan-Animate-2は多種多様なキャラクタや動作パターンにまたがる高忠実なアニメーション結果が得られることが示された。
さらなる研究とコミュニティ開発を促進するため、Wan-Animate-2-Baseモデルウェイトを一般向けに公開します。
関連論文リスト
- LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time [26.186924581106087]
本稿では,リアルタイムストリーミングと安定な長周期生成を10億スケールで組み合わせたアニメーションシステムを提案する。
LiveAnimateは、最初の30秒から最終分まで、知覚品質とアイデンティティをほぼ一定に維持する。
これらの結果は、インタラクティブなフルボディアニメーションのための品質、レイテンシ、持続時間において、新たな操作ポイントを確立する。
論文 参考訳(メタデータ) (2026-08-12T07:35:52Z) - 3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement [51.22216867938034]
映像生成のための再構成3次元環境において,人間の動作軌跡を制御できる3次元モーションアニメーションフレームワークを提案する。
我々は,視点適応型潜伏融合機構を設計し,映像視認性マスキングによる点-雲の幾何前兆を生成過程に注入する。
2つの標準的な人体画像アニメーションベンチマークデータセットによる実験は、関連する映像生成のメカティクスにおける芸術的状況に対する我々の手法の顕著な改善を実証している。
論文 参考訳(メタデータ) (2026-06-29T16:22:54Z) - IM-Animation: An Implicit Motion Representation for Identity-decoupled Character Animation [58.297199313494]
インプシット法は、動画から直接動作の意味をキャプチャするが、動作と外観の絡み合いやアイデンティティの漏洩に悩まされる。
本稿では,フレームごとの動作をコンパクトな1次元モーショントークンに圧縮する新しい暗黙の動作表現を提案する。
本手法では,3段階のトレーニング戦略を用いて,トレーニング効率を高め,高い忠実性を確保する。
論文 参考訳(メタデータ) (2026-02-07T11:17:20Z) - DreamActor-M2: Universal Character Image Animation via Spatiotemporal In-Context Learning [24.808926786222376]
本研究では,DreamActor-M2を提案する。DreamActor-M2は,動作条件をコンテキスト内学習問題として再定義する汎用アニメーションフレームワークである。
まず、参照の出現と動きの手がかりを統一された潜在空間に融合させることにより、入力モダリティギャップを橋渡しする。
次に、擬似的クロスアイデンティティトレーニングペアをキュレートする自己ブートストラップデータ合成パイプラインを導入する。
論文 参考訳(メタデータ) (2026-01-29T13:43:17Z) - Animate-X++: Universal Character Image Animation with Dynamic Backgrounds [32.04255747303296]
Animate-X++は、擬人化文字を含む様々な文字タイプ向けのDiTに基づく普遍的なアニメーションフレームワークである。
動作表現を強化するために,暗黙的かつ明示的な方法で動画から包括的な動作パターンをキャプチャするPose Indicatorを導入する。
第2の課題として、アニメーションとTI2Vタスクを共同でトレーニングするマルチタスクトレーニング戦略を導入する。
論文 参考訳(メタデータ) (2025-08-13T03:11:28Z) - UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation [53.16986875759286]
We present a UniAnimate framework to enable efficient and long-term human video generation。
我々は、姿勢案内やノイズビデオとともに参照画像を共通の特徴空間にマッピングする。
また、ランダムノイズ入力と第1フレーム条件入力をサポートする統一ノイズ入力を提案する。
論文 参考訳(メタデータ) (2024-06-03T10:51:10Z) - Zero-shot High-fidelity and Pose-controllable Character Animation [89.74818983864832]
イメージ・ツー・ビデオ(I2V)生成は、単一の画像からビデオシーケンスを作成することを目的としている。
既存のアプローチは、キャラクターの外観の不整合と細部保存の貧弱さに悩まされている。
文字アニメーションのための新しいゼロショットI2VフレームワークPoseAnimateを提案する。
論文 参考訳(メタデータ) (2024-04-21T14:43:31Z) - AnimateZero: Video Diffusion Models are Zero-Shot Image Animators [63.938509879469024]
我々はAnimateZeroを提案し、事前訓練されたテキスト・ビデオ拡散モデル、すなわちAnimateDiffを提案する。
外観制御のために,テキスト・ツー・イメージ(T2I)生成から中間潜伏子とその特徴を借りる。
時間的制御では、元のT2Vモデルのグローバルな時間的注意を位置補正窓の注意に置き換える。
論文 参考訳(メタデータ) (2023-12-06T13:39:35Z) - MagicAnimate: Temporally Consistent Human Image Animation using
Diffusion Model [74.84435399451573]
本稿では、特定の動きシーケンスに従って、特定の参照アイデンティティのビデオを生成することを目的とした、人間の画像アニメーションタスクについて検討する。
既存のアニメーションは、通常、フレームウォーピング技術を用いて参照画像を目標運動に向けてアニメーションする。
MagicAnimateは,時間的一貫性の向上,参照画像の忠実な保存,アニメーションの忠実性向上を目的とした,拡散に基づくフレームワークである。
論文 参考訳(メタデータ) (2023-11-27T18:32:31Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。