論文の概要: TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation
- arxiv url: http://arxiv.org/abs/2607.24359v1
- Date: Mon, 27 Jul 2026 12:35:32 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-28 22:34:15.418238
- Title: TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation
- Title(参考訳): TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation
- Abstract要約: リアルタイムのロングフォームデジタル・ヒューマン・ジェネレーションは、音声視覚コンテンツを拡張するための因果モデルに依存している。
Methodは、数ステップのジョイントオーディオビデオ生成のためのアンカー誘導永続メモリフレームワークである。
- 参考スコア(独自算出の注目度): 17.429355392513354
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Real-time long-form digital-human generation relies on causal models to extend audio-visual content while preserving subject appearance and audio-video synchronization across successive segments. A bounded cache retains local motion and phonetic context but discards older evidence, whereas attending to the complete generated history is computationally expensive and can propagate accumulated errors. We present \method, an anchor-guided persistent-memory framework for few-step joint audio-video generation. The framework preserves an immutable visual anchor, compresses completed video and audio blocks into fixed-capacity dynamic states, and retrieves those states through modality-specific residual attention without extending the active cache. A reference-aware modulation method additionally conditions video features on dynamic and anchor appearance statistics. Anchor-preserving causal-context distillation varies rollout horizon, prefix provenance, and cache-history reliability while keeping the immutable visual anchor unperturbed. By separating persistent memory from stage-local denoising dependencies, \method further admits stage-parallel execution across blocks, accelerating autoregressive inference without pipeline-specific retraining. We evaluate long-form video continuations with appearance, temporal, synchronization, facial, and speech diagnostics. Results show that \method preserves stable appearance across prompt-conditioned segments and strong audio-visual synchronization under autoregressive generation. Our project page is https://taoliveaigc.github.io/TaoMate.
- Abstract(参考訳): リアルタイム・ロングフォーム・デジタル・ヒューマン・ジェネレーションは、連続したセグメント間の主観的外観とオーディオ・ビデオ同期を保ちながら、音声・視覚的コンテンツを拡張するための因果モデルに依存している。
境界キャッシュは局所的な動きと音韻的文脈を保持するが、古い証拠を捨てる。
数ステップのジョイントオーディオビデオ生成のためのアンカーガイド型永続メモリフレームワークであるShamethodを提案する。
このフレームワークは、不変のビジュアルアンカーを保持し、完了したビデオブロックとオーディオブロックを固定容量の動的状態に圧縮し、アクティブキャッシュを拡張することなく、モダリティ特異的な残留注意によりそれらの状態を取得する。
基準対応変調法は、動的およびアンカーの出現統計にビデオ特徴を付加する。
アンカー保存型因果語蒸留はロールアウト水平線、プレフィックス前駆体、キャッシュ・ヒストリーの信頼性に変化があり、不変の視覚アンカーは飽和しない。
ステージローカルなdenoising依存性から永続メモリを分離することで、ブロック間でのステージ並列実行がさらに許可され、パイプライン固有のリトレーニングなしで自動回帰推論が加速される。
出現・時間・同期・顔・音声診断による長大な映像の連続性の評価を行った。
その結果,<method</method</method</method</method</method</method</method</method</method</method</method</method</method</method</method</method</method</method</method</method</method</method</method</method</method</method</method</method</method
私たちのプロジェクトページはhttps://taoliveaigc.github.io/TaoMateです。
関連論文リスト
- TetherCache: Stabilizing Autoregressive Long-Form Video Generation with Gated Recall and Trusted Alignment [51.33418612284208]
ドリフト耐性長ビデオ生成のためのトレーニングフリーでプラグアンドプレイのキャッシュ管理戦略であるTetherCacheを提案する。
Gated Recall with Attention-Diversity Balancingは、ゲートスコアを使用して長距離メモリフレームを選択する。
TAMEは、信頼されたコンテキスト分布に統計を合わせることで、新しくリコールされたメモリトークンを軽量に編集する。
論文 参考訳(メタデータ) (2026-06-11T08:16:08Z) - FadeMem: Distance-Aware Memory Consolidation for Autoregressive Video Diffusion [59.207505503284715]
FadeMemは、歴史的なKVブロックを固定キャッシュ予算の下で時間階層に整理する。
新しい歴史はきめ細かいエントリとして挿入され、古い隣のエントリは徐々にマージされる。
実験では、既存の有界キャッシュ戦略よりも、被験者の一貫性、背景安定性、時間的コヒーレンスが改善された。
論文 参考訳(メタデータ) (2026-06-09T10:22:18Z) - Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution [25.63670341165374]
ビデオモデルは、証拠が保存されていないときに進化する状態を維持すべきであるが、現在のジェネレータは割り込み時に隠れた状態を凍結することが多い。
本稿では,メモリ指向データ,イベント認識トレーニング,キャッシュ型適応による動的メモリ動作を実現するフレームワークであるReMindを紹介する。
論文 参考訳(メタデータ) (2026-05-25T01:30:41Z) - Echo-Forcing: A Scene Memory Framework for Interactive Long Video Generation [48.476317015122625]
Echo-Forcingは、インタラクティブなロングビデオ生成のためのトレーニング不要のシーンメモリフレームワークである。
キャッシュのバウンダリでスムーズなトランジション、ハードカット、長距離シーンリコールをサポートする。
論文 参考訳(メタデータ) (2026-05-15T14:33:09Z) - CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing [76.74048814837336]
映画ダビングは、ターゲット映像中の唇の動きと同期しながら、参照音声の音声アイデンティティを保持する音声を合成することを目的としている。
既存の方法は正確なリップシンクを達成できず、持続時間レベルでの明示的なアライメントによって自然性を欠いている。
認知同期拡散変換器(CoSync-DiT)により駆動される新しいフローマッチング型フィルムダビングフレームワークを提案する。
論文 参考訳(メタデータ) (2026-04-14T05:03:57Z) - Screen, Match, and Cache: A Training-Free Causality-Consistent Reference Frame Framework for Human Animation [44.20260674331104]
FrameCacheは、Screen、Cache、Matchで構成されるトレーニング不要の3段階フレームワークである。
標準ベンチマークの実験では、FrameCacheは時間的コヒーレンスと視覚的安定性を一貫して改善している。
論文 参考訳(メタデータ) (2025-12-13T08:45:03Z) - Mixture of Contexts for Long Video Generation [72.96361488755986]
我々は長文ビデオ生成を内部情報検索タスクとして再放送する。
本稿では,学習可能なスパークアテンション・ルーティング・モジュールであるMixture of Contexts (MoC) を提案する。
データをスケールしてルーティングを徐々に分散させていくと、そのモデルは計算を適切な履歴に割り当て、アイデンティティ、アクション、シーンを数分のコンテンツで保存する。
論文 参考訳(メタデータ) (2025-08-28T17:57:55Z) - InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing [66.48064661467781]
我々は、アイデンティティ、象徴的なジェスチャー、カメラ軌跡を維持するために参照を戦略的に保存する新しいパラダイムであるスパースフレームビデオダビングを導入する。
無限長長列ダビング用に設計されたストリーミングオーディオ駆動型ジェネレータであるInfiniteTalkを提案する。
HDTF、CelebV-HQ、EMTDデータセットの総合評価は、最先端の性能を示している。
論文 参考訳(メタデータ) (2025-08-19T17:55:23Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。