論文の概要: Dexterous Tactile World Model
- arxiv url: http://arxiv.org/abs/2609.34286v1
- Date: Mon, 28 Sep 2026 04:30:09 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-03 22:12:07.510873
- Title: Dexterous Tactile World Model
- Title(参考訳): Dexterous Tactile World Model
- Abstract要約: 我々は,エゴセントリックな操作の将来のフレーム予測のためのビデオワールドモデルであるDexterous Tactile World Model (DTWM)を提案する。
我々は、ビデオトークン内の対応する手の位置にあるゼロ触覚残差を介して、各手の触覚信号に予めトレーニングされたビデオ拡散変換器を条件付けする。
DTWMは手の動きの過小評価を23%から9%に減らし,手領域の知覚誤差を7.4%減らした。
- 参考スコア(独自算出の注目度): 33.05866318949503
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: World models for manipulation are typically trained from video, yet the events that determine how manipulation unfolds, such as making and releasing contact, are difficult to observe visually and are often easier to sense through touch. We present the Dexterous Tactile World Model (DTWM), a video world model for future-frame prediction of egocentric manipulation from both observed video and tactile signals from a glove worn on each hand. We condition a pretrained video diffusion transformer on each hand's tactile signal through a zero-initialized residual at the corresponding hand location in the video tokens, while a causal mask prevents predicted frames from accessing future information. Compared with a vision-only model matched in architecture, parameters, and training, DTWM reduces the underestimation of hand motion from 23% to 9%, while reducing the perceptual error in the hand region by 7.4% across three training runs per model. The benefit also increases over the prediction horizon, with the improvement in the later predicted chunks being about 4.1x larger than in the first. DTWM also outperforms other visual-tactile world models under the same setting, and training with touch improves future-frame prediction even when no touch is available at inference. Ablations show that the model benefits from both the magnitude and spatial location of force: replacing the tactile signal with binary contact states, either per hand or per location, increases prediction error. The observed course of the force indicates whether the interaction will persist or change.
- Abstract(参考訳): 操作のための世界モデルは典型的にはビデオから訓練されるが、接触の生成や解放といった操作の展開方法を決定するイベントは視覚的に観察することが困難であり、しばしば触覚による理解が容易である。
本稿では,両手に装着した手袋による観察映像と触覚信号の両方から,エゴセントリックな操作を将来予測するためのビデオワールドモデルであるDexterous Tactile World Model (DTWM)を提案する。
我々は、ビデオトークン内の対応する手の位置におけるゼロ初期化残差を介して、各手の触覚信号に事前訓練されたビデオ拡散変換器を条件とし、因果マスクは、予測フレームが将来の情報にアクセスするのを防ぐ。
アーキテクチャ、パラメータ、トレーニングで一致した視覚のみのモデルと比較すると、DTWMは手の動きの過小評価を23%から9%に減らし、手領域の知覚誤差をモデル毎に3回のトレーニング実行で7.4%削減する。
この利点は予測の地平線を超えて増大し、後の予測チャンクの改善は第1の予測よりも約4.1倍大きい。
DTWMは、同じ設定で他の視覚触覚世界モデルよりも優れており、タッチによるトレーニングは、推論時にタッチが利用できない場合でも、将来のフレーム予測を改善する。
アブレーションは、このモデルが力の大きさと空間的位置の両方の利点があることを示し、触覚信号を2つの接触状態に置き換えることで予測誤差を増大させる。
観察された力のコースは、相互作用が持続するかどうかを示す。
関連論文リスト
- DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation [62.08802141030416]
本稿では,指先を個別に符号化するビジュオ触覚WAMであるDexTacWAMを紹介する。
触覚潜伏剤をビデオ拡散世界モデルに注入し, 共同ビジュオ触覚世界モデリングを行う。
論文 参考訳(メタデータ) (2026-09-21T17:55:24Z) - ME-Dex 1.0: Bringing Heterogeneous Tactile Sensing into World Action Modeling [17.299603717233328]
本稿では,共同視覚,触覚,行動学習のための世界行動触覚モデルであるME-Dex-1.0(MachEmbodied-Dex-1.0)を提案する。
多ソースの不均一な触覚入力をサポートするために、カノニカルハンドモデルと統一触覚オートエンコーダマップは、触覚観測を空間空間と潜時空間に共有する。
論文 参考訳(メタデータ) (2026-09-18T08:00:06Z) - TouchSight: Bare-Handed Tactile Prediction from Egocentric Video via Generative Visual Augmentation [7.514180931864231]
我々は,高密度全手接触力予測のための単眼自我中心型視覚フレームワークであるTouchSightを提案する。
TwinTouch-20H: 20時間ペアの視覚データを構築し, 生成的ビデオモデルが手作業で記録した記録を手作業で再現する。
TouchSightは、光沢のあるビデオと生成された裸手ビデオの両方から高密度な力を予測し、OakInk2の以前の接触予測方法より優れています。
論文 参考訳(メタデータ) (2026-09-17T13:58:43Z) - TacWAM: Anchor-Guided World Action Model with Mechanics-Aware Tactile Prediction [15.274487606802124]
World Action Models (WAM) は将来の状態予測とロボットアクション生成を組み合わせるが、既存のアプローチは視覚的未来に大きく依存している。
本稿では、この課題に3つのステップで対処する、メカニック対応の触覚WAMであるTacWAMを紹介する。
筆者らは, 脆弱な把握, 表面接触, 動的手動操作を含む4つの実世界の接触リッチな操作タスクに対して, TacWAMを評価した。
論文 参考訳(メタデータ) (2026-07-30T15:47:01Z) - Tactile-WAM: Touch-Aware World Action Model with Tactile Asymmetric Attention [53.84395207661823]
触覚非対称注意機構(TAAM)を備えたタッチ対応WAMであるTactile-WAMを提案する。
ManiFeelでは、Tactile-WAMは平均成功率を38.9%改善し、コンタクトリッチタスクでは86%改善した。
論文 参考訳(メタデータ) (2026-06-25T06:54:56Z) - VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs [47.982092015932444]
Video-Action Models (VAM) は、インテリジェンスを具現化するための有望なフレームワークとして登場した。
本稿では,触覚を接地信号として組み込んだマルチモーダル世界モデリングフレームワークであるVideo-Tactile Action Model (VTAM)を紹介する。
VTAMは、触覚ストリームでトレーニング済みのビデオトランスフォーマーを軽量なモダリティ転送ファインタニングで強化する。
論文 参考訳(メタデータ) (2026-03-24T17:45:06Z) - OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation [57.133721026727706]
textbfOmniViTacは,16ドルのタスクと100ドル以上のオブジェクトからなる21,000ドル以上のトラジェクトリからなる大規模ビズオタクティルアクションデータセットである。
我々は4つの密結合モジュールを統合する世界モデルベースのビジュオ触覚操作フレームワークである textbf OmniVTA を提案する。
論文 参考訳(メタデータ) (2026-03-19T17:52:42Z) - OPENTOUCH: Bringing Full-Hand Touch to Real-World Interaction [93.88239833545623]
OpenTouchは、最初のインザワイルドなエゴセントリックなフルハンド触覚データセットです。
触覚信号は,理解のためのコンパクトで強力なキューを提供する。
我々は,マルチモーダルな自我中心の知覚,具体的学習,接触に富むロボット操作の促進を目指す。
論文 参考訳(メタデータ) (2025-12-18T18:18:17Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。