論文の概要: Goodbye Drift: Anchored Tree Sampling for Long-Horizon Video-to-Video Generation
- arxiv url: http://arxiv.org/abs/2605.20476v1
- Date: Tue, 19 May 2026 20:40:48 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-05-21 19:19:56.369395
- Title: Goodbye Drift: Anchored Tree Sampling for Long-Horizon Video-to-Video Generation
- Title(参考訳): Goodbye Drift:長距離ビデオ・ビデオ・ジェネレーションのためのアンチョレッドツリーサンプリング
- Authors: Matthew Bendel, Stephen W. Bailey, Mithilesh Vaidya, Sumukh Badam, Xingzhe He,
- Abstract要約: ロングホライゾンビデオ生成は2つの中間問題に悩まされている。
まず、ビデオの品質が時間の経過とともに低下するドリフトがある。
第二に、オブジェクト永続性の問題として現れる連続性の問題があります。
- 参考スコア(独自算出の注目度): 5.125325999976446
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Long-horizon video generation suffers from two intertwined issues. First, there is drift, where video quality degrades over time. Second, there are continuity issues which manifest as object permanence issues, or improperly rendering transient content (e.g., an object that appears in non-consecutive frames changing color/style). Recent work has focused on autoregressive distillation techniques that attack both problems simultaneously. We instead choose to focus on drift directly and introduce \textbf{Anchored Tree Sampling (ATS)}: a training-free inference-time scheduler that replaces left-to-right rollout with sparse-to-dense, anchor-bounded imputation organized as a tree. A root call produces sparse anchors over the full horizon, recursive refinement generates intermediate anchors, and final leaf spans are synthesized between neighboring anchors. This reduces the critical path from $K$ sequential rollout steps to $L+1$ tree-hierarchical steps and converts horizon-compounding drift into anchor-bounded drift. We focus on V2V generation in the \emph{static-camera} regime, where sparse anchors over the horizon are well approximated by the dense conditioning signal, and the base model can produce them without retraining. We evaluate ATS against two contemporary autoregressive baselines on Wan $2.1$ $+$ VACE, across five conditioning modalities (inpainting, outpainting, edge, pose, depth). We show that ATS outperforms both competitors in overall quality, as well as in drift prevention. We additionally demonstrate stable $\geq 40$-minute generation on LTX-$2.3$ across the same five modalities. We conclude by proposing a path forward to extend ATS to arbitrarily long T2V generation, as well as the dynamic-camera and multi-shot regimes.
- Abstract(参考訳): ロングホライゾンビデオ生成は2つの中間問題に悩まされている。
まず、ビデオの品質が時間の経過とともに低下するドリフトがある。
第二に、オブジェクト永続性問題として現れる連続性問題や、過渡的なコンテンツ(例えば、色やスタイルを変える非連続的なフレームに現れるオブジェクト)を不適切にレンダリングする問題があります。
最近の研究は、両方の問題を同時に攻撃する自己回帰蒸留技術に焦点を当てている。
そこで我々は,左から右へのロールアウトに置き換えるトレーニングフリーな推論時スケジューラである‘textbf{Anchored Tree Sampling(ATS)’を導入する。
根の呼び出しは全地平線上でスパースアンカーを生成し、再帰的精製は中間アンカーを生成し、最終葉のスパンは隣接するアンカー間で合成される。
これにより、クリティカルパスは、$K$シーケンシャルなロールアウトステップから$L+1$ツリー階層ステップへと減少し、水平方向のドリフトをアンカーバウンドドリフトに変換する。
地平線上のスパースアンカーを高密度条件付信号でよく近似し, ベースモデルで再学習することなくV2V生成を行う。
We evaluate ATS against two contemporary autoregressive baselines on Wan $2.1$ $$ VACE, across five conditioning modalities (inpainting, outpainting, edge, pose, depth)。
以上の結果から,ATSはドリフト防止だけでなく,総合的な品質の競争相手よりも優れていたことが示唆された。
さらに、安定な$\geq 40$- minutes 生成をLTX-$2.3$で同じ5つのモードで示す。
我々は、ATSを任意の長さのT2V世代に拡張する経路と、ダイナミックカメラとマルチショットレジームを提案して結論付けた。
関連論文リスト
- FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching [68.01498128172214]
本稿では,アーキテクチャに依存しない,追加のトレーニングを必要としない長大なビデオ生成のための,斬新でシンプルな推論時間アプローチを提案する。
提案手法では,隣接するウィンドウからのクリーンサンプルをemphTweedieマッチングでブレンドし,テキストbfmanifoldの制約と重複領域間の時間的一貫性を強制する。
本手法は, 時間的一貫性と視覚的品質において, トレーニング不要, 自己回帰ベースラインを両立させながら, ネイティブウィンドウ長よりも数倍長大のビデオを生成する。
論文 参考訳(メタデータ) (2026-05-20T08:55:37Z) - Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation [50.55853275866995]
リアルタイムのインタラクティブなビデオ生成には、低レイテンシ、ストリーミング、コントロール可能なロールアウトが必要である。
本稿では,フレームワイドの自己回帰を1~2ステップのサンプリングで行うという,よりアグレッシブな設定について検討する。
原理的かつスケーラブルなパイプラインである textbfCausal Forcing++ を提案する。
論文 参考訳(メタデータ) (2026-05-14T17:46:36Z) - A$^2$RD: Agentic Autoregressive Diffusion for Long Video Consistency [49.594617336113714]
A$2$RDは、ビデオセグメント・バイ・セグメンテーションを合成し自己改善する閉ループプロセスとして、長いビデオ合成を定式化する。
i) モダリティ間の動画の進行を追跡するマルチモーダルビデオメモリ、(ii) 自然な進行と視覚的整合性のために生成モード間で切り替える適応セグメント生成、(iii) フレームとビデオレベルの各セグメントを自己改善してエラー伝播を防ぐ階層的テストタイム自己改善の3つのコアコンポーネントから構成される。
論文 参考訳(メタデータ) (2026-05-07T20:35:46Z) - Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning [41.7663185398555]
ViTLは2段階の長ビデオQAフレームワークで、問題関連区間を初期化して固定トークン予算を保存する
ViTLは最大8.6%まで到達し、長時間のQAと時間的グラウンドでは50%少ないフレーム入力を実現している。
論文 参考訳(メタデータ) (2025-10-05T04:03:31Z) - StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text [58.49820807662246]
本稿では,80,240,600,1200以上のフレームをスムーズな遷移で自動回帰的に生成するStreamingT2Vを紹介する。
私たちのコードは、https://github.com/Picsart-AI-Research/StreamingT2V.comで利用可能です。
論文 参考訳(メタデータ) (2024-03-21T18:27:29Z) - Towards Smooth Video Composition [59.134911550142455]
ビデオ生成には、時間とともに動的コンテンツを伴う一貫した永続的なフレームが必要である。
本研究は, 生成的対向ネットワーク(GAN)を用いて, 任意の長さの映像を構成するための時間的関係を, 数フレームから無限までモデル化するものである。
単体画像生成のためのエイリアスフリー操作は、適切に学習された知識とともに、フレーム単位の品質を損なうことなく、スムーズなフレーム遷移をもたらすことを示す。
論文 参考訳(メタデータ) (2022-12-14T18:54:13Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。