論文の概要: DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation
- arxiv url: http://arxiv.org/abs/2609.24976v1
- Date: Mon, 21 Sep 2026 17:55:24 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-22 20:29:01.456146
- Title: DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation
- Title(参考訳): DexTacWAM:Dexterous ManipulationのためのVisuo-Tactile World-Action Model
- Abstract要約: 本稿では,指先を個別に符号化するビジュオ触覚WAMであるDexTacWAMを紹介する。
触覚潜伏剤をビデオ拡散世界モデルに注入し, 共同ビジュオ触覚世界モデリングを行う。
- 参考スコア(独自算出の注目度): 62.08802141030416
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain largely vision-centric and therefore cannot directly model these contact dynamics. We present DexTacWAM, a visuo-tactile WAM that encodes each fingertip independently, aggregates the resulting features through a finger- and pose-aware tactile compressor, and injects the tactile latent into a video diffusion world model for joint visuo-tactile world modeling. Across six contact-rich dexterous manipulation tasks on a 22-DoF bimanual platform, DexTacWAM achieves the highest score on every task, averaging 70.6 versus 38.0 for the strongest baseline. Ablations attribute the gain to modeling contact evolution as part of the predicted world state rather than tactile conditioning alone: removing tactile world modeling reduces the four-task mean from 74.7 to 26.6 while keeping the same tactile features and action expert. After four hours of tactile-encoder adaptation with a frozen pretrained vision VAE, our continual vision-to-touch learning extends the pretrained video model to touch using roughly 100 demonstrations per task without tactile midtraining, while retaining visual prediction quality within 0.5 dB of vision-only counterparts. The compressor retains 89.4% of pre-fusion contact recall while enabling 2.26x faster training and 1.29x faster inference. Together, these results show that pretrained video priors can be extended to distributed multi-finger contact dynamics in a data- and compute-efficient manner.
- Abstract(参考訳): 有害な操作は、視界から部分的にしか観察できない接触力学に依存する。
近年の World-Action Models (WAM) は、アクション生成と予測的ビデオワールドモデリングを結合しているが、主に視覚中心であり、そのためこれらの接触ダイナミクスを直接モデル化することはできない。
本稿では,指先を個別に符号化し,指とポーズを意識した触覚圧縮機を通じて特徴を集約し,その触覚をビデオ拡散世界モデルに注入する。
22-DoF双対プラットフォーム上の6つの接触リッチなデキスタラスな操作タスクの中で、デックスタックワムは最強のベースラインに対して平均70.6対38.0で全てのタスクで最高スコアを達成している。
触覚の世界モデリングを取り除くことで、4つのタスクの平均は74.7から26.6に減少し、同じ触覚の特徴とアクションエキスパートを維持している。
4時間にわたる触覚エンコーダ適応の後、連続的な視覚-触覚学習は、視力のみの0.5dB以内の視覚予測品質を維持しながら、触覚中級トレーニングなしでタスク毎の約100のデモを使用して、トレーニング済みの映像モデルにタッチするように拡張した。
圧縮機は89.4%のプレフュージョンコンタクトリコールを保持し、2.26倍の高速トレーニングと1.29倍の高速推論を可能にした。
これらの結果から,事前学習したビデオは,データと計算効率のよい分散マルチフィンガーコンタクトダイナミックスに拡張可能であることが示された。
関連論文リスト
- STAR: Sparse Tactile Representation Learning in Vision-Tactile-Language-Action Models for Dexterous Manipulation [24.489844571043808]
悪質な操作には、協調した多指制御と効果的な触覚フィードバックが必要である。
ロボットプラットフォームと遠隔操作システムを構築し,200時間の両面的な操作データセットを収集する。
視覚触覚-言語-アクションモデルのための統合学習法STARを提案する。
論文 参考訳(メタデータ) (2026-09-11T07:57:24Z) - VT-WAM: Visual-Tactile World Action Model for Contact-Rich Manipulation [22.0330577788623]
コンタクトリッチな操作には、局所的な変形、圧力、滑り、摩擦に反応するポリシーが必要である。
既存の視覚触覚ポリシーは通常、触覚観察を直接行動予測に供給する。
本稿では,未来の視覚予測,触覚変形予測,行動予測を共同で学習する視覚触覚世界行動モデルであるVT-WAMを紹介する。
論文 参考訳(メタデータ) (2026-07-02T17:58:36Z) - T-Rex: Tactile-Reactive Dexterous Manipulation [139.71263755530654]
本稿では,新しい時相触覚VQ-VAEエンコーダを備えた可変レートMixture-of-Transformers (MoT)アーキテクチャを提案する。
微妙な力制御と変形可能な物体操作を必要とする12の操作課題に対して,触覚反応が有効であることを示す。
論文 参考訳(メタデータ) (2026-06-15T17:59:55Z) - VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs [47.982092015932444]
Video-Action Models (VAM) は、インテリジェンスを具現化するための有望なフレームワークとして登場した。
本稿では,触覚を接地信号として組み込んだマルチモーダル世界モデリングフレームワークであるVideo-Tactile Action Model (VTAM)を紹介する。
VTAMは、触覚ストリームでトレーニング済みのビデオトランスフォーマーを軽量なモダリティ転送ファインタニングで強化する。
論文 参考訳(メタデータ) (2026-03-24T17:45:06Z) - OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation [57.133721026727706]
textbfOmniViTacは,16ドルのタスクと100ドル以上のオブジェクトからなる21,000ドル以上のトラジェクトリからなる大規模ビズオタクティルアクションデータセットである。
我々は4つの密結合モジュールを統合する世界モデルベースのビジュオ触覚操作フレームワークである textbf OmniVTA を提案する。
論文 参考訳(メタデータ) (2026-03-19T17:52:42Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。