論文の概要: Robots Acquire Manipulation Skills in Seconds from a Single Human Video
- arxiv url: http://arxiv.org/abs/2607.20033v2
- Date: Thu, 23 Jul 2026 03:38:42 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-24 14:12:06.751758
- Title: Robots Acquire Manipulation Skills in Seconds from a Single Human Video
- Title(参考訳): ロボットが人間のビデオから秒間にマニピュレーションスキルを取得
- Authors: Guangyan Chen, Meiling Wang, Te Cui, Zichen Zhou, Qi Shao, Shalfun Li, Hang Su, Roy Gan, Hao Wang, Mengyin Fu, Yi Yang, Yufeng Yue,
- Abstract要約: HOST(Human-to-robot One-Shot Skill AcquisiTion)は、ロボットが人間のビデオから数秒でスキルを習得することを可能にするフレームワークである。
HOSTは、自己接地予測のカスケードを通じて、スキル獲得を解決する。
- 参考スコア(独自算出の注目度): 30.999247562623875
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: The ability to acquire skills rapidly and effortlessly while retaining those already mastered is essential for robots. However, current methods still rely on a cumbersome training-time loop that is costly and slow, while eroding skills already mastered. In this paper, we introduce HOST (Human-to-robot One-Shot Skill AcquisiTion), a framework that enables a robot to acquire skills in seconds from a single human video while retaining previously mastered skills. HOST resolves skill acquisition through a cascade of self-grounded prediction. It first estimates the robot's progress within the demonstrated task, then translates the upcoming progression into the robot's own future observations, and finally derives actions from these predicted observations. This cascade is trained on targets coupled to the video demonstration, obtained by mapping the robot trajectory and the video demonstration onto a shared task progress manifold, then redefining each target to align with the future progression of the video. HOST thereby enables the robot to actively follow the demonstrated procedure and adapt it to the robot's embodiment. HOST acquires novel skills at inference time from a single human video in an average of 29 seconds and achieves a 62% average success rate. It exceeds the zero-shot baseline by 45% while retaining previously mastered skills. HOST even exceeds the baseline fine-tuned on 50 robot demonstrations per task while requiring 50 times fewer demonstrations and acquiring each skill 507 times faster. Additional information about HOST is available on the project website.
- Abstract(参考訳): 既に習得しているスキルをロボットに保持しながら、迅速かつ努力的に習得する能力は、ロボットにとって不可欠である。
しかし、現在のメソッドは、コストがかかり、遅いトレーニングタイムループに依存していますが、浸食スキルはすでに熟達しています。
本稿では,ロボットが1つの人間ビデオから数秒でスキルを習得することを可能にするフレームワークであるHOST(Human-to-robot One-Shot Skill AcquisiTion)を紹介する。
HOSTは、自己接地予測のカスケードを通じて、スキル獲得を解決する。
ロボットはまず、実証されたタスクの中でロボットの進捗を推定し、次に次の進行をロボット自身の将来の観察に翻訳し、最終的にこれらの予測された観察から行動を引き出す。
このカスケードは、ロボットの軌道と映像のデモを共有タスク進行多様体にマッピングし、各目標を再定義して、映像の将来の進行に合わせることによって、映像のデモンストレーションに結びついた目標に基づいて訓練される。
これにより、HOSTにより、ロボットは、実証された手順を積極的に追従し、ロボットの体格に適応することができる。
HOSTは、1人の人間のビデオから平均29秒で推論時に新しいスキルを取得し、平均62%の成功率を達成する。
ゼロショットベースラインを45%上回り、それまでのマスタードスキルを維持している。
HOSTは50倍のデモを要し、それぞれのスキルを507倍早く獲得する。
HOSTに関する追加情報はプロジェクトのWebサイトにある。
関連論文リスト
- Robot Self-Improvement via Human-Video Dynamics Models [52.14862052988773]
人間のビデオは、ロボットのエンボディメント間で伝達されるエンボディメントに依存しない動作、ダイナミクス、および値表現を学ぶのに利用できることを示す。
DGAC(Dynamics-Guided Action Correction)は,これらの適応モデルを用いて故障状態の修復を行うトレーニングフリーアプローチである。
以上の結果から,人間の先行とロボットの失敗が組み合わさって,スケーラブルな自律的政策改善を可能にすることが示唆された。
論文 参考訳(メタデータ) (2026-06-19T13:17:27Z) - HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos [58.9564236347451]
HumanEgoは、人間とロボットのエンボディメントギャップを橋渡しするフレームワークである。
それは、人間のデモを、手動オブジェクトの相互作用の実体レベルの表現へと持ち上げる。
HumanEgoは、ロボットのデータフリー、ハードウェア非依存、データ効率、ゼロショットの人間とロボットの転送を可能にする。
論文 参考訳(メタデータ) (2026-05-24T08:26:41Z) - H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos [58.006918399913665]
本稿では,通常の人間と物体のインタラクションビデオからモーション一貫性のあるロボット操作ビデオに変換するビデオ間翻訳フレームワークを提案する。
私たちのアプローチでは、ロボットビデオのセットのみをトレーニングするために、ペアの人間ロボットビデオは必要とせず、システムを拡張しやすくしています。
テスト時にも同じプロセスを人間のビデオに適用し、人間の行動を模倣する高品質なロボットビデオを生成する。
論文 参考訳(メタデータ) (2025-12-10T07:59:45Z) - Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-Training [69.54948297520612]
ジェネラリストの具体化エージェントを学ぶことは、主にアクションラベル付きロボットデータセットの不足に起因して、課題を提起する。
これらの課題に対処するための新しい枠組みを導入し、人間のビデオにおける生成前トレーニングと、少数のアクションラベル付きロボットビデオのポリシー微調整を組み合わせるために、統一された離散拡散を利用する。
提案手法は, 従来の最先端手法と比較して, 高忠実度な今後の計画ビデオを生成し, 細調整されたポリシーを強化する。
論文 参考訳(メタデータ) (2024-02-22T09:48:47Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。