論文の概要: Action Motifs: Self-Supervised Hierarchical Representation of Human Body Movements
- arxiv url: http://arxiv.org/abs/2604.28173v1
- Date: Thu, 30 Apr 2026 17:55:01 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-05-01 16:31:54.239872
- Title: Action Motifs: Self-Supervised Hierarchical Representation of Human Body Movements
- Title(参考訳): 行動モチーフ:人体運動の自己監督的階層的表現
- Authors: Genki Kinoshita, Shu Nakamura, Ryo Kawahara, Shohei Nobuhara, Yasutomo Kawanishi, Ko Nishino,
- Abstract要約: 本稿では,原子間関節運動を捉えるアクション原子と,その時間的構成によって形成されるアクションモチーフからなる階層的表現を提案する。
我々は、ネストされた潜伏トランスフォーマーであるA4Merを導出し、人間のポーズデータから完全に自己教師された方法でこの階層表現を学ぶ。
- 参考スコア(独自算出の注目度): 38.685761162850774
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Effective human behavior modeling requires a representation of the human body movement that capitalizes on its compositionality. We propose a hierarchical representation consisting of Action Atoms that capture the atomic joint movements and Action Motifs that are formed by their temporal compositions and encode similar body movements found across different overall human actions. We derive A4Mer, a nested latent Transformer to learn this hierarchical representation from human pose data in a fully self-supervised manner. A4Mer splits a 3D pose sequence into variable-length segments and represents each segment as a single latent token (Action Atoms). Through bottom-up representation learning, temporal patterns composed of these Action Atoms, which capture meaningful temporal spans of reusable, semantic segments of body movements, naturally emerge (Action Motifs). A4Mer achieves this with a unified pretext task of masked token prediction in their respective latent spaces. We also introduce Action Motif Dataset (AMD), a large-scale dataset of multi-view human behavior videos with full SMPL annotations. We introduce a novel use of cameras by mounting them on the feet to achieve their frame-wise annotations despite frequent and heavy body occlusions. Experimental results demonstrate the effectiveness of A4Mer for extracting meaningful Action Motifs, which significantly benefit human behavior modeling tasks including action recognition, motion prediction, and motion interpolation.
- Abstract(参考訳): 効果的な人間の行動モデリングには、その構成性に重きを置く人間の身体運動の表現が必要である。
本研究では,原子間関節運動をとらえる行動原子と,その時間的構成によって形成される行動モチーフとからなり,人間の行動全体にわたって見られる類似体の動きをエンコードする行動原子からなる階層的表現を提案する。
我々は、ネストされた潜伏トランスフォーマーであるA4Merを導出し、人間のポーズデータから完全に自己教師された方法でこの階層表現を学ぶ。
A4Merは3Dポーズシーケンスを可変長セグメントに分割し、各セグメントを単一の潜在トークン(Action Atoms)として表現する。
ボトムアップ表現学習(ボトムアップ表現学習)を通じて、これらのアクション・アトム(Action Atoms)で構成され、身体運動の再利用可能な意味的なセグメントの有意義な時間的スパンをキャプチャする(Action Motifs)。
A4Merは、それぞれの潜在空間におけるマスク付きトークン予測の統一されたプレテキストタスクでこれを達成している。
我々はまた、フルSMPLアノテーションを備えたマルチビュー人間行動ビデオの大規模データセットであるAction Motif Dataset (AMD)も導入した。
頻繁で重度な身体閉塞にもかかわらず、フレームワイドアノテーションを実現するために、カメラを足に装着することで、新しい利用法を提案する。
A4Merは,行動認識,動作予測,動作補間など,人間の行動モデリングタスクに有意な効果をもたらす。
関連論文リスト
- MotionWeaver: Holistic 4D-Anchored Framework for Multi-Humanoid Image Animation [22.502601281241724]
MotionWeaverは、マルチヒューマノイド画像アニメーションのためのエンドツーエンドフレームワークである。
我々は、同一性に依存しない動きを抽出し、対応する文字に明示的に結合する統合された動き表現を導入する。
また,ビデオラテントで映像表現を融合するために,共有4次元空間を構成する包括的4次元アンコールパラダイムを提案する。
論文 参考訳(メタデータ) (2026-02-11T03:03:44Z) - CoopDiff: Anticipating 3D Human-object Interactions via Contact-consistent Decoupled Diffusion [62.93198247045824]
3Dヒューマンオブジェクトインタラクション(HOI)は,人間の将来の動きとその操作対象を,歴史的文脈で予測することを目的としている。
そこで我々は,人間と物体の運動モデリングを分離するために,2つの異なる分岐を用いた接触非結合拡散フレームワークCoopDiffを提案する。
論文 参考訳(メタデータ) (2025-08-10T03:29:17Z) - Capturing Rich Behavior Representations: A Dynamic Action Semantic-Aware Graph Transformer for Video Captioning [13.411096520754507]
既存のビデオキャプション手法は、単にオブジェクトの振舞いの浅いあるいは単純化した表現を提供するだけである。
本稿では,オブジェクトの振る舞いの本質を包括的に把握する動的アクション意味認識グラフ変換器を提案する。
論文 参考訳(メタデータ) (2025-02-19T14:16:47Z) - Scaling Up Dynamic Human-Scene Interaction Modeling [58.032368564071895]
TRUMANSは、現在利用可能な最も包括的なモーションキャプチャーHSIデータセットである。
人体全体の動きや部分レベルの物体の動きを複雑に捉えます。
本研究では,任意の長さのHSI配列を効率的に生成する拡散型自己回帰モデルを提案する。
論文 参考訳(メタデータ) (2024-03-13T15:45:04Z) - Task-Oriented Human-Object Interactions Generation with Implicit Neural
Representations [61.659439423703155]
TOHO: 命令型ニューラル表現を用いたタスク指向型ヒューマンオブジェクトインタラクション生成
本手法は時間座標のみでパラメータ化される連続運動を生成する。
この研究は、一般的なヒューマン・シーンの相互作用シミュレーションに向けて一歩前進する。
論文 参考訳(メタデータ) (2023-03-23T09:31:56Z) - SportsCap: Monocular 3D Human Motion Capture and Fine-grained
Understanding in Challenging Sports Videos [40.19723456533343]
SportsCap - 3Dの人間の動きを同時に捉え、モノラルな挑戦的なスポーツビデオ入力からきめ細かなアクションを理解するための最初のアプローチを提案する。
本手法は,組込み空間に先立って意味的かつ時間的構造を持つサブモーションを,モーションキャプチャと理解に活用する。
このようなハイブリッドな動き情報に基づいて,マルチストリーム空間時空間グラフ畳み込みネットワーク(ST-GCN)を導入し,詳細なセマンティックアクション特性を予測する。
論文 参考訳(メタデータ) (2021-04-23T07:52:03Z) - Action2Motion: Conditioned Generation of 3D Human Motions [28.031644518303075]
我々は3Dで人間の動作シーケンスを生成することを目的としている。
それぞれのサンプル配列は、自然界の体動力学に忠実に類似している。
新しい3DモーションデータセットであるHumanAct12も構築されている。
論文 参考訳(メタデータ) (2020-07-30T05:29:59Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。