論文の概要: UniDex-ViTac: Learning Unified Visuo-Tactile Dexterous Manipulation Policy from Human Video Data
- arxiv url: http://arxiv.org/abs/2609.16504v1
- Date: Tue, 15 Sep 2026 01:52:33 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-16 14:56:08.300727
- Title: UniDex-ViTac: Learning Unified Visuo-Tactile Dexterous Manipulation Policy from Human Video Data
- Title(参考訳): UniDex-ViTac:人間のビデオデータからUnified Visuo-Tactile Dexterous Manipulation Policyを学ぶ
- Abstract要約: We present UniDex-ViTac, a framework that using human-video-guided Simulation to generate robot demonstrations with fingertip contact observed。
我々は1万の擬似軌道を収集し、1つのAction Chunking with Transformers (ACT)ベースのジェネラリストを訓練する。
接触拡大された構成は、ポイントクラウドのみのベースラインで55.5%に対して、シミュレーションで68.3%のマクロ平均成功を実現している。
- 参考スコア(独自算出の注目度): 0.45880283710344055
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Human videos provide demonstrations of dexterous manipulation but lack robot-executable actions and tactile measurements. We present UniDex-ViTac, a framework that uses human-video-guided simulation to generate robot demonstrations paired with fingertip contact observations for training a deployable visuo-tactile policy. Object-specific residual reinforcement learning specialists adapt annotated human-object interaction references to a robotic arm-hand system. Their successful rollouts pair final robot action targets with robot-side fingertip contact observations. From 50 human demonstrations across ten objects, we collect 10,000 simulated trajectories to train a single Action Chunking with Transformers (ACT) based generalist. The policy combines point clouds, proprioception, and four binary contact signals encoded through fingertip labels and a separate token, without requiring human references or privileged object identity and pose at deployment. The contact-augmented configuration achieves 68.3% macro-average success in simulation, compared with 55.5% for the point-cloud-only baseline. Without real-robot demonstrations or policy fine-tuning, it succeeds in 73/110 physical trials (66.4%) across six seen and five unseen objects, compared with 60/110 (54.5%) for the baseline, an increase of 11.8 percentage points. These results support the feasibility of learning a unified visuo-tactile dexterous manipulation policy from video-guided simulated interactions. Project page: https://unidex-vitac.github.io/
- Abstract(参考訳): 人間のビデオは、巧妙な操作のデモを提供するが、ロボットの実行可能な動作と触覚測定が欠如している。
We present UniDex-ViTac, a framework that using human-video-guided Simulation to generate robot demos with paired with fingertip contact observed for training a deploymentable visuo-tactile policy。
オブジェクト固有の残留強化学習スペシャリストは、ロボットアームハンドシステムへの注釈付き人間とオブジェクトのインタラクション参照を適応する。
彼らの成功したロールアウトは、ロボット側の指先接触観察と最後のロボットアクションターゲットのペアである。
10個のオブジェクトにまたがる50の人間デモから、シミュレーションされた1万の軌道を集め、1つのAction Chunking with Transformers (ACT)ベースのジェネラリストを訓練します。
このポリシーは、ポイントクラウド、プロプレセプション、および指先ラベルと別のトークンで符号化された4つのバイナリコンタクト信号を組み合わせて、人間の参照や特権オブジェクトのアイデンティティを必要とせず、デプロイ時にポーズを取る。
接触拡大された構成は、ポイントクラウドのみのベースラインで55.5%に対して、シミュレーションで68.3%のマクロ平均成功を実現している。
実際のロボットのデモンストレーションやポリシーの微調整がなければ、6つの見えない物体と5つの見えない物体の73/110(66.4%)で成功し、ベースラインの60/110(54.5%)は11.8ポイント上昇した。
これらの結果は、ビデオ誘導型シミュレートされたインタラクションから、統一的なビズータクタクタクタブルな操作ポリシーを学習する可能性を支持する。
プロジェクトページ: https://unidex-vitac.github.io/
関連論文リスト
- Weave: Learning Whole-Body Dexterous Loco-Manipulation from Human-Object Interactions [6.968853416384178]
ヒューマノイドと物体の相互作用を学習するには、ロボットと物体の動きの両方を制御するために、全身のバランス、移動、器用な接触を調整する必要がある。
我々は、捕獲された人間の実演から全身のデキスタス・ヒューマノイド・オブジェクトの相互作用を学習するための統合フレームワークWeaveを提案する。
論文 参考訳(メタデータ) (2026-09-15T06:04:36Z) - Dex-X: Learning Visual-Tactile Dexterous Manipulation From Human Videos with Simulated Interaction [30.46515328565822]
シミュレーションにより人間の映像から視覚触覚のデキスタラスな操作を学習するためのフレームワークであるDEX-Xを提案する。
そこで,本研究では,触覚の制御を行う物理接地力学を用いて,手動物体の相互作用を再現する。
多様な把握と接触に富んだツール使用タスクにまたがるデクスタラスハンドアームプラットフォーム上でのゼロショットsim-to-real転送を実証する。
論文 参考訳(メタデータ) (2026-09-07T16:47:39Z) - Robot Learning from Human Demonstrations: Handwritten Alphabet Trajectories and Human-Likeness Evaluation [0.26763498831034044]
本稿では,人間のようなロボットの動きを実演から学習するための枠組みを提案する。
参加者22名から3,142名の手書きによる実演のデータセットを収集した。
21名の被験者によるユーザスタディは、生成された軌跡の人間的類似性を評価した。
論文 参考訳(メタデータ) (2026-08-06T16:12:18Z) - Dexterous Point Policy: Learning Point-based Dexterous Hand Policies from Human Demonstrations [47.549615826832856]
Dexterous Point Policyは、人間のビデオから直接繊細な操作ポリシーを学び、ロボットのデモを必要としない。
課題関連オブジェクトと人間の手の3Dキーポイントを生ビデオから抽出し,これらのキーポイント上で自己回帰変換器を訓練する。
デクサラスポイント・ポリシーは、ピック・アンド・プレイスとツール・ユースにまたがるリアル・ロボットの一連のタスクで75.0%の成功を収め、最先端のVLAベースラインは1.0%にしか達していない。
論文 参考訳(メタデータ) (2026-06-09T09:13:36Z) - Human-to-Robot Interaction: Learning from Video Demonstration for Robot Imitation [5.967530183571141]
人間とロボットの模倣学習パイプラインは、ロボットが非構造化ビデオデモから直接操作スキルを取得することを可能にする。
鍵となる革新は、学習プロセスを2つの異なる段階に分離するモジュラーフレームワークである。
ロボット操作では,全ての動作の平均成功率は87.5%であり,タスク達成で100%,複雑なピック・アンド・プレイス操作で90%に達する。
論文 参考訳(メタデータ) (2026-02-22T13:26:27Z) - DexImit: Learning Bimanual Dexterous Manipulation from Monocular Human Videos [56.64773686434068]
DexImitは、人間の操作映像を物理的に妥当なロボットデータに変換する自動フレームワークである。
DexImitは、インターネットまたはビデオ生成モデルから、人間のビデオに基づいて大規模なロボットデータを生成することができる。
ツールの使用、長距離タスク、きめ細かい操作を含む多様な操作タスクを処理できる。
論文 参考訳(メタデータ) (2026-02-10T18:59:02Z) - Dexterous Manipulation Policies from RGB Human Videos via 3D Hand-Object Trajectory Reconstruction [24.49384094440561]
我々は,RGBのヒューマンビデオから直接デクスタラスな操作を学習する,デバイスフリーのフレームワークであるVIDEOMANIPを提案する。
シミュレーションでは、学習した把握モデルはインスパイアハンドを用いて20種類のオブジェクトに対して70.25%の成功率を達成する。
実世界では、RGBビデオから訓練された操作ポリシーは、LEAPハンドを使用して7つのタスクで平均62.86%の成功率を達成する。
論文 参考訳(メタデータ) (2026-02-09T18:56:02Z) - HumanX: Toward Agile and Generalizable Humanoid Interaction Skills from Human Videos [74.43500240121476]
我々はHumanXについて紹介する。HumanXは人間の動画を、ヒューマノイドのための汎用的で現実的なインタラクションスキルにコンパイルするフルスタックのフレームワークである。
HumanXは、従来の方法より8倍高い一般化成功を達成する。
論文 参考訳(メタデータ) (2026-02-02T18:53:01Z) - MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training [102.850162490626]
人間のロボットによる相互模倣事前学習による視覚-言語-行動モデルであるMiVLAを提案する。
MiVLAは、最先端のVLAよりも優れた、強力な改良された一般化能力を実現する。
論文 参考訳(メタデータ) (2025-12-17T12:59:41Z) - Feel the Force: Contact-Driven Learning from Humans [52.36160086934298]
操作中のきめ細かい力の制御は、ロボット工学における中核的な課題である。
We present FeelTheForce, a robot learning system that model human tactile behavior to learn force-sensitive control。
提案手法は,5つの力覚的操作タスクで77%の成功率を達成した,スケーラブルな人間の監督において,堅牢な低レベル力制御を実現する。
論文 参考訳(メタデータ) (2025-06-02T17:57:52Z) - Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning [3.9738951919572827]
本稿では,Voxelized RGB-D空間におけるロボットデモを用いて,RGBビデオから人間デモを明示的にモデル化するフレームワークを提案する。
本稿では,人間の意図モデリングのためのResNetベースの視覚符号化と,ボクセルに基づくロボット行動予測のためのPerceiver Transformerを組み合わせる。
論文 参考訳(メタデータ) (2025-04-14T21:14:51Z) - Learning Manipulation by Predicting Interaction [85.57297574510507]
本稿では,インタラクションを予測して操作を学習する一般的な事前学習パイプラインを提案する。
実験の結果,MPIは従来のロボットプラットフォームと比較して10%から64%向上していることがわかった。
論文 参考訳(メタデータ) (2024-06-01T13:28:31Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。