論文の概要: Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation
- arxiv url: http://arxiv.org/abs/2609.34182v2
- Date: Tue, 29 Sep 2026 02:47:09 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-30 21:28:46.71879
- Title: Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation
- Title(参考訳): Dexterous Manipulationのための人間のデモからの一元的視覚触覚-反応モデリング
- Abstract要約: 有害な操作には触覚フィードバックが必要である。
人間のデモは、多様な触覚インタラクションのよりスケーラブルなソースを提供する。
我々は、人間の触覚データを活用して、巧妙な操作ポリシーを改善します。
- 参考スコア(独自算出の注目度): 19.452443308926487
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Dexterous manipulation requires tactile feedback. However, robot tactile demonstrations are difficult to scale,because dexterous-hand teleoperation provides limited tactile feedback to the operator. In contrast, human demonstrations offer a substantially more scalable source of diverse tactile interactions. Motivated by a simple premise: hands can change, but the underlying physics of interaction does not. We leverage human tactile data to improve dexterous manipulation policies. Specifically, we first build a tactile motion-capture system that synchronously records images, tactile signals, and hand motions. Using this system, we construct the UVTA dataset spanning five contact-rich tasks, with 1,000 human demonstrations covering diverse interaction patterns and 150 robot demonstrations per task. To transfer the underlying physics of human interaction to robot control, we propose a Unified Visual-Tactile-Action Model that maps both embodiments into aligned tactile and action representations and jointly predicts future action and tactile trajectories. The joint objective enables human demonstrations to supervise contact-aware representation learning, while only robot actions are executed during deployment. In real-robot evaluations across five tasks, our method achieves an average success rate of 70%, outperforming the strongest visual-tactile baseline, which achieves 29%, and an architecture ablation, which achieves 42%. Performance improves consistently with additional human demonstrations and exhibits no saturation at 1,000 demonstrations per task, validating the effectiveness of scalable human tactile data for dexterous manipulation. Project page is available at https://uni-vta.github.io/.
- Abstract(参考訳): 有害な操作には触覚フィードバックが必要である。
しかし, 遠隔操作は操作者に限られた触覚フィードバックを与えるため, ロボットの触覚実証はスケールが困難である。
対照的に、人間のデモは多様な触覚相互作用のかなりスケーラブルなソースを提供する。
単純な前提によって動機付けられる:手は変化しうるが、基礎となる相互作用の物理学は変化しない。
我々は、人間の触覚データを活用して、巧妙な操作ポリシーを改善します。
具体的には、まず、画像、触覚信号、手の動きを同期的に記録する触覚モーションキャプチャシステムを構築します。
本システムを用いて,5つのタスクにまたがるUVTAデータセットを構築する。
人間のインタラクションの基礎となる物理をロボット制御に伝達するために,両実施形態を協調した触覚と行動表現にマッピングし,将来の行動と触覚の軌跡を共同で予測する統一視覚触覚-行動モデルを提案する。
共同目的により、人間による実演では、ロボットアクションのみがデプロイ中に実行される一方で、接触認識型表現学習を監督することができる。
5タスクにわたる実ロボット評価において,本手法は平均成功率70%を達成し,29%の視覚触覚ベースライン,42%のアーキテクチャアブレーションを達成している。
パフォーマンスは、追加の人間のデモンストレーションと一貫して改善され、タスク毎の1000のデモでは飽和が示されず、拡張性のある人間の触覚データの有効性を検証する。
プロジェクトページはhttps://uni-vta.github.io/.com/で公開されている。
関連論文リスト
- Skel-WAM: A Hand-Skeleton-Conditioned World Action Model for Human-to-Robot Manipulation Transfer [61.50662046794264]
Skel-WAMは,ハンドスケルトン・モーション・インタフェースの統一による差異をブリッジする世界アクション・モデルである。
主な洞察は、人間の動きとロボットの動きを共通のハンドトポロジーで調整し、シーン内の地面の動きを、明示的な手の動きをエンコードする2.5Dキーポイントと組み合わせることである。
ビデオとキーポイントの専門家は、Mixture-of-Transformersを通じて視覚と骨格のダイナミクスを共同で学習し、ロボット訓練されたアクションエキスパートはこれらの予測を実行可能なコントロールにマップする。
論文 参考訳(メタデータ) (2026-09-18T09:06:15Z) - Dex-X: Learning Visual-Tactile Dexterous Manipulation From Human Videos with Simulated Interaction [30.46515328565822]
シミュレーションにより人間の映像から視覚触覚のデキスタラスな操作を学習するためのフレームワークであるDEX-Xを提案する。
そこで,本研究では,触覚の制御を行う物理接地力学を用いて,手動物体の相互作用を再現する。
多様な把握と接触に富んだツール使用タスクにまたがるデクスタラスハンドアームプラットフォーム上でのゼロショットsim-to-real転送を実証する。
論文 参考訳(メタデータ) (2026-09-07T16:47:39Z) - AdvDex: Learning Dexterous Manipulation from Human Demonstrations via Joint-Aligned Actions and Adversarial Learning [25.241504281004243]
本稿では,人間とロボットのデモから,器用な操作を学ぶためのフレームワークであるAdvDexを紹介する。
まず、人間の操作デモの大規模マルチモーダルデータセットであるOmniShareを紹介する。
第2に,手首ポーズが$mathrmSE(3)$および15指関節からなる標準動作表現であるJAAS(Joint-Aligned Action Space)を提案する。
第3に,学習した視覚表現における具体的情報を減らすために,ドメイン・アドバイサル・ラーニングを用いる。
論文 参考訳(メタデータ) (2026-08-14T07:19:06Z) - ActiveGlasses: Learning Manipulation with Active Vision from Ego-centric Human Demonstration [51.69384671495837]
ActiveGlassesは、エゴ中心の人間のデモからロボット操作を学習するシステムである。
スマートグラスに装着されたステレオカメラは、データ収集とポリシー推論の両方のための唯一の認識装置として機能する。
ゼロ・トランスファーを可能にするために,デモからオブジェクト・トラジェクトリを抽出し,オブジェクト中心のポイント・クラウド・ポリシーを用いて操作と頭部運動を協調的に予測する。
論文 参考訳(メタデータ) (2026-04-09T17:59:08Z) - EgoHumanoid: Unlocking In-the-Wild Loco-Manipulation with Robot-Free Egocentric Demonstration [67.13034606664333]
EgoHumanoidは、エゴセントリックな人間のデモを使って視覚言語アクションポリシーを共同訓練する最初のフレームワークである。
スケーラブルな人的データ収集のためのポータブルシステムを開発した。
論文 参考訳(メタデータ) (2026-02-10T18:59:03Z) - Feel the Force: Contact-Driven Learning from Humans [52.36160086934298]
操作中のきめ細かい力の制御は、ロボット工学における中核的な課題である。
We present FeelTheForce, a robot learning system that model human tactile behavior to learn force-sensitive control。
提案手法は,5つの力覚的操作タスクで77%の成功率を達成した,スケーラブルな人間の監督において,堅牢な低レベル力制御を実現する。
論文 参考訳(メタデータ) (2025-06-02T17:57:52Z) - RealDex: Towards Human-like Grasping for Robotic Dexterous Hand [64.33746404551343]
本稿では,人間の行動パターンを取り入れた手の動きを正確に把握する先駆的データセットであるRealDexを紹介する。
RealDexは、現実のシナリオにおける認識、認識、操作を自動化するためのヒューマノイドロボットを進化させる上で、大きな可能性を秘めている。
論文 参考訳(メタデータ) (2024-02-21T14:59:46Z) - Robot Synesthesia: In-Hand Manipulation with Visuotactile Sensing [15.970078821894758]
視覚的・触覚的な感覚入力を活用して手動操作を可能にするシステムを提案する。
ロボット・シンセシス(Robot Synesthesia)は、人間の触覚と視覚の合成にインスパイアされた、新しい点の雲に基づく触覚表現である。
論文 参考訳(メタデータ) (2023-12-04T12:35:43Z) - Visual Imitation Made Easy [102.36509665008732]
本稿では,ロボットへのデータ転送を容易にしながら,データ収集プロセスを単純化する,模倣のための代替インターフェースを提案する。
我々は、データ収集装置やロボットのエンドエフェクターとして、市販のリーチ・グラブラー補助具を使用する。
我々は,非包括的プッシュと包括的積み重ねという2つの課題について実験的に評価した。
論文 参考訳(メタデータ) (2020-08-11T17:58:50Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。