論文の概要: Efficient Audio-Visual Generation via Synchrony-Aware Cross-Modal Sparse Attention
- arxiv url: http://arxiv.org/abs/2608.15522v1
- Date: Sun, 16 Aug 2026 04:26:05 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-18 19:59:03.327602
- Title: Efficient Audio-Visual Generation via Synchrony-Aware Cross-Modal Sparse Attention
- Title(参考訳): 同期型クロスモーダルスパースアテンションによる高能率オーディオ映像生成
- Authors: Shengchuan Gao, Teng Hu, Bohao Feng, Luchen Li, Wenqiang Wang, Hongqian Deng, Ran Yi,
- Abstract要約: 効率的な音声視覚生成のための同期対応アクセラレーションフレームワークを提案する。
提案手法は,映像品質,音声品質,音声とビデオの同期を保ちながら,推論効率を向上させる。
- 参考スコア(独自算出の注目度): 31.052048543229464
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Recent audio-visual generation models can synthesize synchronized video and sound in a unified diffusion process, but their inference cost remains high because long video token sequences require repeated attention computation across denoising steps.A variety of acceleration techniques have been developed for video generation models, including low-bit quantization, attention sparsification, and feature caching.However, since these methods are originally designed for video generation, directly applying them to audio-visual models overlooks the interactions between the audio and video branches and may therefore disrupt audio-video synchronization.We present a synchronization-aware acceleration framework for efficient audio-visual generation.Our key observation is that bidirectional audio-video cross-attention reveals structured interactions between the two branches, with high responses often concentrated on a few sound-related visual and temporal regions.Guided by this interaction pattern, we introduce a protected sparse attention strategy that preserves high-fidelity computation for synchronization-critical tokens while sparsifying redundant attention interactions.By explicitly accounting for cross-modal dependence during acceleration, our method improves inference efficiency while keeping video quality, audio quality, and audio-video synchronization.
- Abstract(参考訳): 近年の音声-視覚生成モデルは、統合拡散過程において、同期映像と音声を合成することができるが、長いビデオトークンシーケンスは、復調ステップ間で繰り返し注意計算を必要とするため、その推論コストが高い。この手法は、低ビット量子化、アテンションスパリフィケーション、特徴キャッシングなど、ビデオ生成モデル向けに様々な加速技術が開発されているが、これらの手法は、もともとビデオ生成用に設計されており、オーディオ-視覚モデルを直接適用することで、オーディオ/映像ブランチ間の相互作用を見越して、オーディオ-視覚同期を阻害する可能性がある。
関連論文リスト
- Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory [20.92714383117306]
Rippleは、モード間リカレントメモリ機構を備えたリアルタイムのジョイントオーディオビデオ生成システムである。
480P解像度で28FPSを実現し、教師よりも高速で、一貫性のあるロングフォーム生成を実現している。
論文 参考訳(メタデータ) (2026-07-29T12:13:01Z) - Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation [50.411841997631484]
We present Unison, a unified framework that promote coherence across the motion, speech, and sound modalities。
We show that Unison achieves state-of-the-art performance in audio perceptual quality and cross-modal synchro。
論文 参考訳(メタデータ) (2026-05-09T06:32:54Z) - UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions [34.27531187147479]
UniAVGenは、ジョイントオーディオとビデオ生成のための統一されたフレームワークである。
UniAVGenは、オーディオオーディオ同期、音色、感情の一貫性において全体的なアドバンテージを提供する。
論文 参考訳(メタデータ) (2025-11-05T10:06:51Z) - Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers [19.226787997122987]
380x640の解像度、24fpsのビデオが多様な音声入力と同期するSyncphonyを提案する。
提案手法は,事前学習したビデオバックボーン上に構築され,同期性を改善するために2つの重要なコンポーネントが組み込まれている。
AVSync15とThe Greatest Hitsデータセットの実験では、Syncphonyは同期精度と視覚的品質の両方で既存のメソッドよりも優れています。
論文 参考訳(メタデータ) (2025-09-26T05:30:06Z) - READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head Generation [55.58089937219475]
本稿では,最初のリアルタイム拡散変換器を用いた音声ヘッド生成フレームワークREADを提案する。
提案手法はまず,VAEを用いて高度に圧縮されたビデオ潜時空間を学習し,音声生成におけるトークン数を大幅に削減する。
また,READは,実行時間を大幅に短縮した競合する音声ヘッドビデオを生成することにより,最先端の手法よりも優れていることを示す。
論文 参考訳(メタデータ) (2025-08-05T13:57:03Z) - MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation [21.216297567167036]
MirrorMeは、LTXビデオモデル上に構築されたリアルタイムで制御可能なフレームワークである。
MirrorMeは映像を空間的に時間的に圧縮し、効率的な遅延空間をデノイングする。
EMTDベンチマークの実験では、MirrorMeの忠実さ、リップシンク精度、時間的安定性が実証されている。
論文 参考訳(メタデータ) (2025-06-27T09:57:23Z) - Live2Diff: Live Stream Translation via Uni-directional Attention in Video Diffusion Models [64.2445487645478]
大規模言語モデルは、テキストやオーディオなどのストリーミングデータの生成において顕著な効果を示している。
本稿では,一方向の時間的注意を向けたビデオ拡散モデルを設計するための最初の試みであるLive2Diffを紹介する。
論文 参考訳(メタデータ) (2024-07-11T17:34:51Z) - Synchformer: Efficient Synchronization from Sparse Cues [100.89656994681934]
コントリビューションには、新しい音声-視覚同期モデル、同期モデルからの抽出を分離するトレーニングが含まれる。
このアプローチは、濃密な設定とスパース設定の両方において最先端の性能を実現する。
また,100万スケールの 'in-the-wild' データセットに同期モデルのトレーニングを拡張し,解釈可能性に対するエビデンス属性技術を調査し,同期モデルの新たな機能であるオーディオ-視覚同期性について検討する。
論文 参考訳(メタデータ) (2024-01-29T18:59:55Z) - Sparse in Space and Time: Audio-visual Synchronisation with Trainable
Selectors [103.21152156339484]
本研究の目的は,一般映像の「野生」音声・視覚同期である。
我々は4つのコントリビューションを行う: (i) スパース同期信号に必要な長時間の時間的シーケンスを処理するために、'セレクタ'を利用するマルチモーダルトランスモデルを設計する。
音声やビデオに使用される圧縮コーデックから生じるアーティファクトを識別し、トレーニングにおいてオーディオ視覚モデルを用いて、同期タスクを人工的に解くことができる。
論文 参考訳(メタデータ) (2022-10-13T14:25:37Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。