論文の概要: InterCMDM: Block-Causal Diffusion for Autoregressive Human Interaction Generation
- arxiv url: http://arxiv.org/abs/2607.01743v1
- Date: Thu, 02 Jul 2026 05:58:15 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-03 19:45:08.689945
- Title: InterCMDM: Block-Causal Diffusion for Autoregressive Human Interaction Generation
- Title(参考訳): InterCMDM: 自己回帰型ヒューマンインタラクション生成のためのブロック因果拡散
- Authors: Qing Yu, Kent Fujiwara,
- Abstract要約: 自己回帰的2人インタラクション生成のためのブロック因果拡散フレームワークを提案する。
InterCMDMはInterHumanとInter-Xの最先端性能を実現し、テキスト・モーションアライメント、リアリズム、長期連続性を改善する。
- 参考スコア(独自算出の注目度): 21.26677650704416
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Text-conditioned human interaction generation must capture both long-range temporal causality within each individual and tightly coupled coordination between partners. Existing interaction diffusion models typically denoise full sequences using bidirectional attention, which obscures causality and hinders streaming and long-horizon generation. Autoregressive alternatives enforce causality but often suffer from temporal drift, leading to coordination degradation and unstable interaction dynamics over time. We propose InterCMDM, a block-causal latent diffusion framework for autoregressive two-person interaction generation. InterCMDM introduces a Dual-Stream Causal Diffusion Transformer that maintains separate causal streams for each person while modeling inter-person dependencies via unified dual-stream attention with multi-task attention masks. These masks unify interaction modeling within a single attention mechanism and support diverse coordination behaviors, including simultaneous actions, reactive responses, leader-follower dynamics, and independent motion. By training a single model across these mask configurations as a form of data augmentation, InterCMDM enables controllable interaction generation by simply selecting the desired attention mask at inference time. Finally, a block-wise diffusion objective enables stable latent rollout over long sequences without repeated decode-encode cycles. InterCMDM achieves state-of-the-art performance on InterHuman and Inter-X, improving text-motion alignment, realism, and long-horizon continuity.
- Abstract(参考訳): テキスト条件付きヒューマンインタラクション生成は、各個人内の長距離時間因果関係と、パートナー間の緊密に結合した調整の両方をキャプチャしなければなりません。
既存の相互作用拡散モデルは通常、双方向の注意を用いて全シーケンスをノイズ化し、因果関係を曖昧にし、ストリーミングやロングホライゾン生成を妨げる。
自己回帰的な代替手段は因果関係を強制するが、しばしば一時的なドリフトに悩まされ、時間とともに協調劣化と不安定な相互作用のダイナミクスを引き起こす。
自動回帰2人インタラクション生成のためのブロック因果拡散フレームワークであるInterCMDMを提案する。
InterCMDMはDual-Stream Causal Diffusion Transformerを導入し、個人ごとに個別の因果ストリームを維持しながら、マルチタスクのアテンションマスクによる統合されたデュアルストリームアテンションを通じて、個人間の依存関係をモデル化する。
これらのマスクは、単一の注意機構内での相互作用モデリングを統一し、同時行動、反応反応、リーダー-フォロワーダイナミクス、独立運動を含む多様な協調行動をサポートする。
これらのマスク構成の1つのモデルをデータ拡張の形でトレーニングすることにより、InterCMDMは、推論時に所望の注目マスクを選択することで、制御可能なインタラクション生成を可能にする。
最後に、ブロックワイド拡散目標により、復号-符号化サイクルを繰り返すことなく、長いシーケンス上で安定した潜時ロールアウトが可能となる。
InterCMDMはInterHumanとInter-Xの最先端性能を実現し、テキスト・モーションアライメント、リアリズム、長期連続性を改善する。
関連論文リスト
- Interact2Ar: Full-Body Human-Human Interaction Generation via Autoregressive Diffusion Models [80.28579390566298]
テキスト条件付き自己回帰拡散モデルであるInteract2Arを導入する。
ハンドキネマティクスは専用のパラレルブランチを通じて組み込まれ、高忠実度フルボディ生成を可能にする。
我々のモデルは、時間的動きの合成、外乱へのリアルタイム適応、ディヤディックからマルチパーソンシナリオへの拡張など、一連のダウンストリームアプリケーションを可能にする。
論文 参考訳(メタデータ) (2025-12-22T18:59:50Z) - Diffusion Forcing for Multi-Agent Interaction Sequence Modeling [52.769202433667125]
MAGNetはマルチエージェントモーション生成のための統合された自己回帰拡散フレームワークである。
フレキシブルな条件付けとサンプリングを通じて、幅広いインタラクションタスクをサポートする。
緊密に同期された活動と、ゆるやかに構造化された社会的相互作用の両方をキャプチャする。
論文 参考訳(メタデータ) (2025-12-19T18:59:02Z) - InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction Graphs [72.5651722107621]
InterAgentはテキスト駆動型物理ベースのマルチエージェントヒューマノイド制御のためのエンドツーエンドフレームワークである。
本稿では,マルチストリームブロックを備えた自己回帰拡散トランスフォーマーを提案する。
また,空間依存性の微粒化を明示的に捉えた対話グラフのエクスセプション表現を提案する。
論文 参考訳(メタデータ) (2025-12-08T10:46:01Z) - Persistent-Transient Duality: A Multi-mechanism Approach for Modeling
Human-Object Interaction [58.67761673662716]
人間は高度に適応可能で、異なるタスク、状況、状況を扱うために異なるモードを素早く切り替える。
人間と物体の相互作用(HOI)において、これらのモードは、(1)活動全体に対する大規模な一貫した計画、(2)タイムラインに沿って開始・終了する小規模の子どもの対話的行動の2つのメカニズムに起因していると考えられる。
本研究は、人間の動作を協調的に制御する2つの同時メカニズムをモデル化することを提案する。
論文 参考訳(メタデータ) (2023-07-24T12:21:33Z) - InterGen: Diffusion-based Multi-human Motion Generation under Complex Interactions [49.097973114627344]
動作拡散プロセスに人間と人間の相互作用を組み込んだ効果的な拡散ベースアプローチであるInterGenを提案する。
我々はまず、InterHumanという名前のマルチモーダルデータセットをコントリビュートする。これは、様々な2人インタラクションのための約107Mフレームで構成され、正確な骨格運動と23,337の自然言語記述を持つ。
本稿では,世界規模での2人のパフォーマーのグローバルな関係を明示的に定式化した対話拡散モデルにおける動作入力の表現を提案する。
論文 参考訳(メタデータ) (2023-04-12T08:12:29Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。