論文の概要: Bridging Semantic and Kinematic Conditions with Diffusion-based Discrete Motion Tokenizer
- arxiv url: http://arxiv.org/abs/2603.19227v1
- Date: Thu, 19 Mar 2026 17:59:51 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-03-20 17:19:06.334212
- Title: Bridging Semantic and Kinematic Conditions with Diffusion-based Discrete Motion Tokenizer
- Title(参考訳): 拡散型離散運動トケナイザを用いたブリッジ・セマンティック・キネマティック条件
- Authors: Chenyang Gu, Mingyuan Zhang, Haozhe Xie, Zhongang Cai, Lei Yang, Ziwei Liu,
- Abstract要約: MoTokは、セマンティックな抽象化をきめ細かな再構築から切り離す離散モーショントークンである。
また,HumanML3Dでは,トークンの6分の1しか使用せず,MaskControl上での制御性と忠実度を大幅に向上する。
- 参考スコア(独自算出の注目度): 55.9892973179428
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Prior motion generation largely follows two paradigms: continuous diffusion models that excel at kinematic control, and discrete token-based generators that are effective for semantic conditioning. To combine their strengths, we propose a three-stage framework comprising condition feature extraction (Perception), discrete token generation (Planning), and diffusion-based motion synthesis (Control). Central to this framework is MoTok, a diffusion-based discrete motion tokenizer that decouples semantic abstraction from fine-grained reconstruction by delegating motion recovery to a diffusion decoder, enabling compact single-layer tokens while preserving motion fidelity. For kinematic conditions, coarse constraints guide token generation during planning, while fine-grained constraints are enforced during control through diffusion-based optimization. This design prevents kinematic details from disrupting semantic token planning. On HumanML3D, our method significantly improves controllability and fidelity over MaskControl while using only one-sixth of the tokens, reducing trajectory error from 0.72 cm to 0.08 cm and FID from 0.083 to 0.029. Unlike prior methods that degrade under stronger kinematic constraints, ours improves fidelity, reducing FID from 0.033 to 0.014.
- Abstract(参考訳): 先行運動生成は、運動制御に優れた連続拡散モデルと、セマンティック・コンディショニングに有効な離散トークンベースのジェネレータの2つのパラダイムに大きく従っている。
本研究では,条件特徴抽出(パーセプション),離散トークン生成(プランニング),拡散に基づく動き合成(コントロル)を含む3段階のフレームワークを提案する。
このフレームワークの中心となるMoTokは、拡散デコーダに運動回復を委譲することで、微細な再構成から意味論的抽象化を分離し、動きの忠実さを保ちながら、コンパクトな単一層トークンを可能にする。
キネマティックな条件では、粗い制約は計画中のトークン生成を導くが、細かい制約は拡散に基づく最適化によって制御中に強制される。
この設計は、キネマティックな詳細がセマンティックトークンの計画を乱すのを防ぐ。
また,HumanML3Dでは,トークンの6分の1しか使用せず,MaskControlの制御性や忠実度を大幅に向上させ,軌道誤差を0.72cmから0.08cm,FIDを0.083から0.029に低減した。
より強いキネマティック制約の下で劣化する従来の方法とは異なり、我々の手法は忠実度を改善し、FIDは0.033から0.014に減少する。
関連論文リスト
- Causal Autoregressive Diffusion Language Model [70.7353007255797]
CARDは厳密な因果注意マスク内の拡散過程を再構成し、単一の前方通過で密集した1対1の監視を可能にする。
我々の結果は,CARDが並列生成のレイテンシの利点を解放しつつ,ARMレベルのデータ効率を実現することを示す。
論文 参考訳(メタデータ) (2026-01-29T17:38:29Z) - Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank [65.00301565190824]
mnameは、外部エンコーダを必要としない、プラグアンドプレイのトレーニングフレームワークである。
mnameは400kのステップでtextbf2.40 の最先端 FID を達成し、同等のメソッドを著しく上回っている。
論文 参考訳(メタデータ) (2025-12-09T14:39:26Z) - $\mathcal{E}_0$: Enhancing Generalization and Fine-Grained Control in VLA Models via Continuized Discrete Diffusion [65.77755100137728]
本稿では、量子化されたアクショントークンを反復的にデノケーションするアクション生成を定式化する、連続的な離散拡散フレームワークであるE0を紹介する。
E0は14の多様な環境において最先端のパフォーマンスを達成し、平均して10.7%強のベースラインを達成している。
論文 参考訳(メタデータ) (2025-11-26T16:14:20Z) - From Tokens to Nodes: Semantic-Guided Motion Control for Dynamic 3D Gaussian Splatting [26.57713792657793]
制御密度と動きの複雑さを一致させる動き適応フレームワークを提案する。
既存の最先端手法に比べて,復元品質と効率が大幅に向上したことを示す。
論文 参考訳(メタデータ) (2025-10-03T05:33:58Z) - Continuously Augmented Discrete Diffusion model for Categorical Generative Modeling [87.34677262370924]
標準離散拡散モデルは、吸収[MASK]トークンにそれらをマッピングすることで、すべての観測されていない状態を同一に扱う。
これは'インフォメーション・ヴォイド'を生成します。そこでは、偽のトークンから推測できるセマンティック情報は、デノイングステップの間に失われます。
連続的拡張離散拡散(Continuously Augmented Discrete Diffusion)は、連続的な潜在空間における対拡散で離散状態空間を拡大するフレームワークである。
論文 参考訳(メタデータ) (2025-10-01T18:00:56Z) - DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding [29.643549839940025]
本稿では、DisCoRD: Rectified Flow Decodingによる連続運動への離散トークンの導入について紹介する。
私たちの中核となる考え方は、条件生成タスクとしてトークンデコーディングをフレーム化することです。
DisCoRDは、HumanML3Dで0.032、KIT-MLで0.169、最先端のパフォーマンスを実現している。
論文 参考訳(メタデータ) (2024-11-29T07:54:56Z) - MaskControl: Spatio-Temporal Control for Masked Motion Synthesis [46.20712809159041]
生成マスク運動モデルに制御性を導入するための最初のアプローチであるMaskControlを提案する。
まず、textitLogits Regularizerは、トレーニング時に暗黙的にロジットを摂り、モーショントークンの分布を制御された関節位置と整列させる。
第2に、textitLogit最適化は、生成した動きを制御された関節位置と正確に一致させるトークン分布を明示的に再設定する。
論文 参考訳(メタデータ) (2024-10-14T17:50:27Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。