論文の概要: KAI: A Kinematic-Aware Interface for Data-Efficient Articulated Object Manipulation
- arxiv url: http://arxiv.org/abs/2607.24493v1
- Date: Mon, 27 Jul 2026 14:29:47 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-28 22:34:15.467735
- Title: KAI: A Kinematic-Aware Interface for Data-Efficient Articulated Object Manipulation
- Title(参考訳): KAI:データ効率の良いArticulated Object Manipulationのためのキネマティック・アウェアインタフェース
- Authors: Yaping Li, Zhaxizhuoma, Qiaojun Yu, Jia Zeng, Dahua Lin, Jiangmiao Pang,
- Abstract要約: 人工物体の操作には、ロボットのデモンストレーションだけでは学べない運動構造を理解する必要がある。
本稿では,音節オブジェクトの運動構造をキャプチャする中間表現であるKAI(Kinematic-Aware Articulation Interface)を紹介する。
解釈可能な幾何学的および運動論的事前を政策学習に組み込むことで、KAIは調音運動の基盤構造に沿った強い帰納的バイアスを与える。
- 参考スコア(独自算出の注目度): 63.408378267309786
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Articulated object manipulation requires an understanding of kinematic structure that is difficult and costly to learn from robot demonstrations alone. We introduce the Kinematic-Aware Articulation Interface (KAI), a structured intermediate representation that captures the kinematic structure of articulated objects. By embedding interpretable geometric and kinematic priors into policy learning, KAI provides a strong inductive bias aligned with the underlying structure of articulated motion. This design effectively improves sample efficiency, with gains particularly pronounced in low-data regimes: across six simulation tasks, our method achieves an average success rate of 82.9%, matching or surpassing baseline performance while using only half the demonstration data. Our method also exhibits robust generalization to unseen backgrounds and visual distractors, transferring from a single clean training environment to cluttered real-world scenes. KAI's action-agnostic design further enables co-training with human interaction videos to enhance real-world robustness: under diverse visual distractions, our method with video co-training achieves over 70% average success rate.
- Abstract(参考訳): 人工物体の操作には、ロボットのデモンストレーションだけで学ぶのが難しくてコストがかかる運動構造を理解する必要がある。
本稿では,音節オブジェクトの運動構造をキャプチャする中間表現であるKAI(Kinematic-Aware Articulation Interface)を紹介する。
解釈可能な幾何学的および運動論的事前を政策学習に組み込むことで、KAIは調音運動の基盤構造に沿った強い帰納的バイアスを与える。
この設計はサンプル効率を効果的に改善し、6つのシミュレーションタスクで平均成功率82.9%を達成し、実演データの半分しか使用せず、ベースライン性能を達成している。
また,本手法は,1つのクリーンなトレーニング環境から散らばった現実世界のシーンへ移動することで,未知の背景や視覚的障害への堅牢な一般化を示す。
KAIのアクション非依存設計により、人間のインタラクションビデオとのコトレーニングにより、実世界のロバスト性を高めることができる。
関連論文リスト
- ConLA: Contrastive Latent Action Learning from Human Videos for Robotic Manipulation [27.54751123419347]
人間のビデオからロボットポリシーを学習するための教師なし事前学習フレームワークであるConLAを提案する。
人間のビデオのみに事前学習を行うことで、実際のロボット軌道事前学習で得られた性能を初めて上回ります。
論文 参考訳(メタデータ) (2026-01-31T06:40:57Z) - Physical Autoregressive Model for Robotic Manipulation without Action Pretraining [65.8971623698511]
我々は、自己回帰ビデオ生成モデルを構築し、物理自己回帰モデル(PAR)を提案する。
PARは、アクション事前トレーニングを必要とせず、物理力学を理解するために、ビデオ事前トレーニングに埋め込まれた世界の知識を活用する。
ManiSkillベンチマークの実験は、PARがPushCubeタスクで100%の成功率を達成したことを示している。
論文 参考訳(メタデータ) (2025-08-13T13:54:51Z) - MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning [66.53533434848369]
密集した表現を学習する動き誘導型自己学習フレームワークを提案する。
6つの画像およびビデオデータセットと4つの評価ベンチマークにおいて、最先端を1%から6%改善する。
論文 参考訳(メタデータ) (2025-06-10T11:20:32Z) - Kinematic-aware Prompting for Generalizable Articulated Object
Manipulation with LLMs [53.66070434419739]
汎用的なオブジェクト操作は、ホームアシストロボットにとって不可欠である。
本稿では,物体のキネマティックな知識を持つ大規模言語モデルに対して,低レベル動作経路を生成するキネマティック・アウェア・プロンプト・フレームワークを提案する。
我々のフレームワークは8つのカテゴリで従来の手法よりも優れており、8つの未確認対象カテゴリに対して強力なゼロショット能力を示している。
論文 参考訳(メタデータ) (2023-11-06T03:26:41Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。