論文の概要: Spatial Grafting: Grounding 3D Features for Flow-Matching Robot Policies
- arxiv url: http://arxiv.org/abs/2609.35249v1
- Date: Mon, 28 Sep 2026 14:17:31 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-02 07:14:26.529821
- Title: Spatial Grafting: Grounding 3D Features for Flow-Matching Robot Policies
- Title(参考訳): 空間グラフト:フローマッチングロボットのためのグラウンドティング3次元特徴
- Abstract要約: そこで本研究では,凍結した再構成特徴をロボット相対幾何学に結合する,汎用的で軽量な空間モジュールを提案する。
空間グラフティングは、メトリックグラウンド化された空間トークンを構築し、クロスアテンションを通じてフローマッチングアクションエキスパートに注入する。
比較した幾何認識ポリシーよりも広範に評価する。
- 参考スコア(独自算出の注目度): 9.75334900626813
- License: http://creativecommons.org/licenses/by-nc-sa/4.0/
- Abstract: Pretrained robot manipulation policies such as vision-language-action models (VLAs) or world-action models (WAMs) leave interaction-relevant metric geometry implicit. Recent breakthroughs in spatial reconstruction can supply the necessary geometry reliably, but their features describe local shape without stating where it lies with respect to the robot. How best to deliver these features to a pretrained policy remains unresolved. We propose Spatial Grafting, a versatile, lightweight spatial module that binds frozen reconstruction features to metric, robot-relative geometry. Spatial Grafting constructs metric-grounded spatial tokens and injects them into the flow-matching action expert through cross-attention, without modifying the host's perceptual pathway, so the host retains the full benefit of its pretraining. We evaluate it more broadly than any geometry-aware policy we compare against: one graft architecture, with no per-host redesign, on two VLAs and two WAMs, across four simulation benchmarks that span short-horizon manipulation, visual robustness, clutter and long-horizon mobile manipulation, and on three real-robot platforms with single- and dual-arm configurations. On RoboTwin 2.0, a dual-arm manipulation benchmark, the graft improves every host across VLAs and WAMs. Grafted $π_{0.5}$ gains 11.3% and 15.6% on clean and randomized scenes, reaching 94.0% and 92.4%, above the strongest published 3D-conditioned policy, WAM4D (93.8% and 89.9%). The margin widens as the horizon lengthens: on tasks from BEHAVIOR-1K, a dual-arm mobile manipulation challenge scored by average task progress, it surpasses the 2025 challenge winner on five of six tasks,by up to 0.47 Q-score, and exceeds a map-conditioned spatial policy on average across the three tasks both report.
- Abstract(参考訳): 視覚言語行動モデル (VLA) や世界行動モデル (WAM) のような事前訓練されたロボット操作ポリシーは、相互作用関連計量幾何学を暗黙的に残す。
空間再構成の最近の進歩は、必要な地形を確実に提供できるが、それらの特徴は、ロボットがどこにあるかを記述することなく、局所的な形状を記述している。
これらの機能を事前訓練されたポリシーにどのように配置するかは、未解決のままだ。
本研究では,凍結した再構成特徴をロボット相対幾何学に結合する,汎用的で軽量な空間モジュールである空間グラフティングを提案する。
空間グラフティング(Spatial Grafting)は、メトリックグラウンド化された空間トークンを構築し、ホストの知覚経路を変更することなく、クロスアテンションを通じてフローマッチングアクションエキスパートに注入する。
2つのVLAと2つのWAM、短い水平操作、視覚的堅牢性、クラッタと長距離移動操作、シングルアームとデュアルアーム構成の3つの実ロボットプラットフォームにまたがる4つのシミュレーションベンチマーク。
デュアルアーム操作ベンチマークであるRoboTwin 2.0では、移植はVLAとWAMのすべてのホストを改善している。
π_{0.5}$ Grafted $π_{0.5}$ gains 11.3% and 15.6% on clean and randomized scenes, up 94.0% and 92.4%, on the highest published 3D-conditioned policy, WAM4D (93.8% and 89.9%)。
BEHAVIOR-1Kのタスクでは、平均的なタスクの進捗によって、2025年のチャレンジの勝者を最大0.47Qスコアで上回り、どちらのレポートでも平均3つのタスクでマップ条件の空間ポリシーを上回ります。
関連論文リスト
- AFUN: Towards an Affordance Foundation Model for Functionality Understanding [12.890216832485647]
我々は,機能理解のための手頃な基礎モデルに向けたステップとして,我々のモデルを提示する。
我々は、異種ロボット、人間、シミュレーション、現実世界のスキャンデータを共有価格スキーマに変換する大規模な標準化データパイプラインを構築します。
私たちのモデルは、4つのベンチマークから8つのテストセットにまたがる大きなマージンで、すべてのベースラインを上回ります。
論文 参考訳(メタデータ) (2026-06-01T17:50:16Z) - PanoEnv: Exploring 3D Spatial Intelligence in Panoramic Environments with Reinforcement Learning [5.308328605042682]
360パノラマ画像は、バーチャルリアリティー、自律運転、総合的なシーン理解のためのロボティクスでますます利用されている。
現在の視覚言語モデル(VLM)は、幾何学的歪みと限定的な3次元監督のため、等角射影(ERP)画像の空間的推論に苦慮している。
合成3D環境から構築した大規模VQAベンチマークであるPanoEnvを紹介する。
我々の7Bモデルは、新しい最先端性能を実現し、全体的な精度を52.93%(+3.59%)、オープンエンド精度を14.83%に改善し、構造化タスク性能を維持した。
論文 参考訳(メタデータ) (2026-02-25T15:12:17Z) - RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics [53.053660003572965]
本稿では,3次元空間参照と計測の両方を初めて実現した3D対応VLMであるRoboTracerを提案する。
RoboTracerは、強化微調整により、多段階のメートル法推論を進める。
本稿では,空間的トレーシングを評価する上で困難なベンチマークであるTraceSpatial-Benchを提案する。
論文 参考訳(メタデータ) (2025-12-15T18:52:43Z) - SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation [63.48859753472547]
SpaceActorは、意味論と幾何学を明確に分離する堅牢なロボット操作のためのフレームワークである。
RLBenchの87.4%で最先端のパフォーマンスを達成し、ノイズの異なる条件下では13.9%から19.4%改善している。
論文 参考訳(メタデータ) (2025-11-12T18:59:08Z) - InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy [138.89177083578213]
空間接地とロボット制御のための統合フレームワークであるInternVLA-M1を紹介する。
InternVLA-M1は、(i)2.3M以上の空間的推論データに基づく空間的グラウンドトレーニングと(ii)空間的に誘導された後トレーニングという、2段階のパイプラインを使用する。
結果: InternVLA-M1 は SimplerEnv Google Robot で+14.6%、WidowX で+17%、LIBERO Franka で+4.3% で、空間誘導なしでその変種を上回った。
論文 参考訳(メタデータ) (2025-10-15T17:30:05Z) - RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics [67.11221574129937]
空間参照は、3D物理世界と相互作用するエンボディロボットの基本的な能力である。
本稿では,まず空間的理解を正確に行うことのできる3次元VLMであるRoboReferを提案する。
RoboReferは、強化微調整による一般化された多段階空間推論を推進している。
論文 参考訳(メタデータ) (2025-06-04T17:59:27Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。