論文の概要: MUGEN: A Unified Framework for Efficient Motion Understanding and Generation
- arxiv url: http://arxiv.org/abs/2607.27581v1
- Date: Thu, 30 Jul 2026 01:55:51 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-31 21:37:00.362293
- Title: MUGEN: A Unified Framework for Efficient Motion Understanding and Generation
- Title(参考訳): MUGEN: 効率的な動作理解と生成のための統一フレームワーク
- Authors: Zhankai Ye, Yukai Jin, Bingyang Wei, Bofan Li, Yusen Wu, Fangyi Li, Shangqian Gao, Xin Liu,
- Abstract要約: コストがかからない一貫した動き言語フレームワークを提案する。
単一の適応長オートエンコーダは、任意の長さの運動をいくつかの連続した潜在スロットに圧縮する。
MUGENは、HumanML3D上のFIDに基づく言語モデルベースラインを導き、リアルモーション参照よりも精度の高い検索を行う。
- 参考スコア(独自算出の注目度): 17.452895169020817
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Grounding human motion in language, and language in motion, is a central step toward physical AI systems that can understand, generate, and communicate human behavior. Unified motion--language systems first coupled the two directions through a shared discrete motion codebook, but quantization limits generation quality. The strongest generators buy quality back at growing cost: stacked residual codebooks enlarge the representation; masked decoding stages, long autoregressive rollouts, and denoising chains of tens to hundreds of steps stretch inference; even the continuous-latent designs among them reach their latent only through an iterative diffusion head; and none of this decoding machinery serves understanding. We therefore propose MUGEN, a unified motion--language framework that pays neither cost: no codebook, one draw. A single adaptive-length autoencoder compresses any-length motion into a few continuous latent slots, the system's only motion representation: the language model generates them for text-to-motion and reads them back for motion understanding. Depth-routed hidden states let each slot read from the transformer depth it needs, and a calibrated head predicts a joint distribution over the full latent set, so a single draw carries the text-conditional, cross-slot variation a description permits. At a decoding cost of K language-model steps, one draw, and one decoder pass, MUGEN leads language-model baselines on FID on HumanML3D while raising retrieval precision above the real-motion reference under the standard evaluator, achieves the best CIDEr and BLEU@4 scores, and surpasses the discrete-token state of the art on every retrieval and alignment metric on SnapMoGen.
- Abstract(参考訳): 言語における人間の動きと動作中の言語は、人間の振る舞いを理解し、生成し、伝達する物理AIシステムへの中心的なステップである。
統一モーション言語システムはまず、共有された離散モーションコードブックを通じて2つの方向を結合するが、量子化は生成品質を制限する。
積み重ねられた残余のコードブックは表現を拡大し、マスクされたデコードステージ、長い自己回帰的なロールアウト、数十から数百ステップのデノイングチェーンは推論を延ばす。
そこで我々は,コストがかからない統合モーション言語フレームワークMUGENを提案する。
単一の適応長オートエンコーダは、任意の長さの動作をいくつかの連続した潜在スロットに圧縮する。
深さが減った隠れ状態は、各スロットを変換器の深さから読み取ることができ、キャリブレーションされたヘッドは完全な潜在集合上の結合分布を予測する。
MUGENは、HumanML3D上のFIDに基づく言語モデルベースラインをリードすると同時に、標準評価器のリアルモーション基準よりも高い検索精度を上げ、最高のCIDErとBLEU@4スコアを獲得し、SnapMoGen上の検索およびアライメントメトリックのそれぞれにおいて、個別の最先端を越えながら、K言語モデルステップのデコードコスト、ワン・ドロー、ワン・デコーダパスのデコードコストにおいて、言語モデルベースラインをリードする。
関連論文リスト
- CLAW: Composable Language-Annotated Whole-body Motion Generation [55.99805728566105]
CLAWは,言語を付加した全身運動データをスケーラブルに生成するためのパイプラインである。
CLAWは運動プランナーから運動プリミティブを構成し、動き、方向、速度、骨盤の高さ、持続時間によってパラメータ化される。
低レベルコントローラは、これらの参照を MuJoCo シミュレーションで追跡し、物理的に接地された軌道を生成する。
論文 参考訳(メタデータ) (2026-04-13T10:02:04Z) - SignSparK: Efficient Multilingual Sign Language Production via Sparse Keyframe Learning [54.232148007248874]
現在の手話生産(SLP)フレームワークは、まさにトレードオフに直面している。
本研究では,スペースを利用した新たなトレーニングパラダイムを提案し,人間の署名の真の基盤となる分布を捉える。
これらの離散的なアンカーから高密度な動きを予測することにより、流体の調音を確実にしながら、回帰から平均への移動を緩和する。
論文 参考訳(メタデータ) (2026-03-11T06:02:36Z) - LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens [19.167250154665812]
LLaMoは、モダリティ固有のMixture-of-Transformersアーキテクチャを通じて、事前訓練された大規模言語モデルを拡張するフレームワークである。
人間の動きを因果連続潜伏空間にエンコードし、デコーダのみのバックボーンで次のトーケン予測パラダイムを維持する。
実験により,LLaMoは一般的な設定で高忠実なテキスト・ツー・モーション生成とモーション・トゥ・テキストキャプションを実現することが示された。
論文 参考訳(メタデータ) (2026-02-12T20:02:21Z) - Advanced Sign Language Video Generation with Compressed and Quantized Multi-Condition Tokenization [20.063863466319326]
SignViPは、複数のきめ細かい条件を組み込んだ新しいフレームワークである。
SignViPは、ビデオ品質の時間的コヒーレンスやセマンティクスの忠実さなど、メトリクス間の最先端のパフォーマンスを実現している。
論文 参考訳(メタデータ) (2025-06-19T02:56:06Z) - MAG: Multi-Modal Aligned Autoregressive Co-Speech Gesture Generation without Vector Quantization [8.605691647343065]
既存の手法では,ジェスチャ生成にベクトル量子化トークンを伴って自己回帰モデルを用いるのが一般的である。
我々は、離散トークン化に頼ることなく、高品質で多様な音声合成のための新しいマルチモーダルアライメントフレームワークMAGを提案する。
論文 参考訳(メタデータ) (2025-03-18T09:02:02Z) - Seamless Human Motion Composition with Blended Positional Encodings [38.85158088021282]
後処理や冗長な復調ステップを伴わずにシームレスなヒューマン・モーション・コンポジション(HMC)を生成する最初の拡散モデルであるフローMDMを紹介する。
我々はBabelとHumanML3Dデータセットの精度、リアリズム、スムーズさの観点から最先端の結果を得る。
論文 参考訳(メタデータ) (2024-02-23T18:59:40Z) - Kosmos-G: Generating Images in Context with Multimodal Large Language Models [117.0259361818715]
現在の被写体駆動画像生成法では、テストタイムチューニングが必要であり、インターリーブされたマルチイメージとテキスト入力を受け付けない。
本稿では,マルチモーダル大規模言語モデルの高度なマルチモーダル認識機能を活用するモデルであるKosmos-Gを提案する。
Kosmos-Gは、インターリーブされたマルチイメージとテキスト入力によるゼロショットの主観的生成の印象的な能力を示す。
論文 参考訳(メタデータ) (2023-10-04T17:28:44Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。