論文の概要: ArtHOI: Taming Foundation Models for Monocular 4D Reconstruction of Hand-Articulated-Object Interactions
- arxiv url: http://arxiv.org/abs/2603.25791v1
- Date: Thu, 26 Mar 2026 18:00:17 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-03-30 21:49:48.219836
- Title: ArtHOI: Taming Foundation Models for Monocular 4D Reconstruction of Hand-Articulated-Object Interactions
- Title(参考訳): ArtHOI:手動物体相互作用の単眼的4次元再構成のためのモデリング基礎モデル
- Authors: Zikai Wang, Zhilu Zhang, Yiqing Wang, Hui Li, Wangmeng Zuo,
- Abstract要約: ArtHOIは最適化ベースのフレームワークで、複数の基礎モデルから事前を統合および洗練する。
特に、オブジェクトのメートル法スケールを最適化するために、適応サンプリング精細法(ASR)を導入する。
また,Multimodal Large Language Model (MLLM) を用いた手オブジェクトアライメント手法を提案する。
- 参考スコア(独自算出の注目度): 48.84720445548848
- License: http://creativecommons.org/licenses/by-nc-sa/4.0/
- Abstract: Existing hand-object interactions (HOI) methods are largely limited to rigid objects, while 4D reconstruction methods of articulated objects generally require pre-scanning the object or even multi-view videos. It remains an unexplored but significant challenge to reconstruct 4D human-articulated-object interactions from a single monocular RGB video. Fortunately, recent advancements in foundation models present a new opportunity to address this highly ill-posed problem. To this end, we introduce ArtHOI, an optimization-based framework that integrates and refines priors from multiple foundation models. Our key contribution is a suite of novel methodologies designed to resolve the inherent inaccuracies and physical unreality of these priors. In particular, we introduce an Adaptive Sampling Refinement (ASR) method to optimize object's metric scale and pose for grounding its normalized mesh in world space. Furthermore, we propose a Multimodal Large Language Model (MLLM) guided hand-object alignment method, utilizing contact reasoning information as constraints of hand-object mesh composition optimization. To facilitate a comprehensive evaluation, we also contribute two new datasets, ArtHOI-RGBD and ArtHOI-Wild. Extensive experiments validate the robustness and effectiveness of our ArtHOI across diverse objects and interactions. Project: https://arthoi-reconstruction.github.io.
- Abstract(参考訳): 既存の手動物体相互作用法(HOI)は、主に剛体物体に限られるが、明瞭な物体の4D再構成法は、通常、対象物や多視点ビデオの事前スキャンを必要とする。
単一の単眼のRGBビデオから4Dの人間と物体の相互作用を再構築することは、まだ解明されていないが重要な課題である。
幸いなことに、最近のファンデーションモデルの進歩は、この非常に不適切な問題に対処する新たな機会を提供する。
この目的のために、最適化ベースのフレームワークであるArtHOIを紹介した。
我々の重要な貢献は、これらの先行の固有の不正確さと物理的不道徳を解決するために設計された、新しい方法論の集合である。
特に、オブジェクトのメートル法スケールを最適化し、その正規化メッシュを世界空間でグラウンド化するためのアダプティブサンプリング精細法(ASR)を導入する。
さらに,マルチモーダル大規模言語モデル (MLLM) を用いた接触推論情報を手動物体メッシュ合成最適化の制約として利用した手動物体アライメント手法を提案する。
包括的評価を容易にするため,ArtHOI-RGBDとArtHOI-Wildという2つの新しいデータセットをコントリビュートする。
広範囲にわたる実験は、さまざまな物体や相互作用におけるArtHOIの堅牢性と有効性を検証する。
プロジェクト: https://arthoi-reconstruction.github.io
関連論文リスト
- ArtLLM: Generating Articulated Assets via 3D LLM [19.814132638278547]
ArtLLMは、完全な3Dメッシュから直接高品質な調音資産を生成するための新しいフレームワークである。
コアとなるのは,大規模な調音データセットに基づいてトレーニングされた,3Dマルチモーダルな大規模言語モデルだ。
実験の結果,ArtLLMは部品配置精度と接合予測の両方で最先端の手法を著しく上回ることがわかった。
論文 参考訳(メタデータ) (2026-03-01T15:07:46Z) - ArtGS: Building Interactable Replicas of Complex Articulated Objects via Gaussian Splatting [66.29782808719301]
コンピュータビジョンにおいて、音声で表現されたオブジェクトを構築することが重要な課題である。
既存のメソッドは、しばしば異なるオブジェクト状態間で効果的に情報を統合できない。
3次元ガウスを柔軟かつ効率的な表現として活用する新しいアプローチであるArtGSを紹介する。
論文 参考訳(メタデータ) (2025-02-26T10:25:32Z) - EasyHOI: Unleashing the Power of Large Models for Reconstructing Hand-Object Interactions in the Wild [79.71523320368388]
本研究の目的は,手動物体のインタラクションを単一視点画像から再構築することである。
まず、手ポーズとオブジェクト形状を推定する新しいパイプラインを設計する。
最初の再構築では、事前に誘導された最適化方式を採用する。
論文 参考訳(メタデータ) (2024-11-21T16:33:35Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。