論文の概要: RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation
- arxiv url: http://arxiv.org/abs/2607.18709v1
- Date: Tue, 21 Jul 2026 05:05:01 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-22 19:05:05.311435
- Title: RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation
- Title(参考訳): RoboInter1.5: 身体的世界モデリングとロボットマニピュレーションのためのホロスティックな中間表現スイート
- Authors: Ziqin Wang, Hao Li, Weijun Wang, Junhao Cai, Jia Zeng, Yilun Chen, Jiangmiao Pang, Si Liu,
- Abstract要約: 本稿では,ロボット操作と擬似世界モデリングの中間表現を拡張し,包括的に表現するRoboInter1.5について紹介する。
RoboInter1.5は、厳密な操作指向の中間表現を中心としたデータ、ベンチマーク、モデルの統一されたリソースを提供する。
- 参考スコア(独自算出の注目度): 46.68319840674789
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Existing robot datasets remain expensive to curate, embodiment-specific, and insufficiently annotated with the fine-grained structure required for generalizable reasoning, execution, or long-horizon environment dynamics simulation. Building on our prior work, RoboInter1.0, we present RoboInter1.5, an extended and holistic suite of intermediate representations for both robotic manipulation and embodied world modeling. RoboInter1.5 provides a unified resource of data, benchmarks, and models centered on dense manipulation-oriented intermediate representations. Specifically, RoboInter-Data contains over 230k manipulation episodes across 571 scenes with dense per-frame annotations covering more than ten types of intermediate representations, including subtasks, primitive skills, object and gripper grounding, segmentation, affordance, grasp poses, contact points, motion traces, etc. Built upon these annotations, RoboInter-VQA introduces spatial and temporal embodied VQA tasks to benchmark and improve the intermediate-representation reasoning capabilities of our RoboInter-VLM. RoboInter-VLA further studies how such representations benefit action execution through implicit, explicit, and modular plan-then-execute paradigms. To better model the physical world, we further introduce RoboInter-World, which leverages intermediate representations as structured conditioning signals for controllable prediction of future world states. Extensive evaluations demonstrate that RoboInter1.5 provides a unified spatiotemporal scaffolding for intermediate representations. Rather than treating intermediate representations merely as interpretable signals, RoboInter1.5 conceptualizes them as a bidirectional interface that both regularizes low-level action spaces and constrains the latent rollouts of open-world physical simulators.
- Abstract(参考訳): 既存のロボットデータセットは、一般的な推論、実行、長期の環境力学シミュレーションに必要なきめ細かい構造で、キュレートし、具体化し、十分に注釈を付けるのに高価である。
ロボット操作と擬似世界モデリングの両方のための中間表現を拡張し包括的に表現するRoboInter1.5を,我々の以前の研究であるRoboInter1.0に基づいて紹介する。
RoboInter1.5は、厳密な操作指向の中間表現を中心としたデータ、ベンチマーク、モデルの統一されたリソースを提供する。
具体的には、RoboInter-Dataには、サブタスク、プリミティブスキル、オブジェクトとグリップグラウンド、セグメンテーション、セグメンテーション、グリップポーズ、コンタクトポイント、モーショントレースなどを含む10種類以上の中間表現を含む、フレーム単位のアノテーションを備えた571シーンにわたる230kの操作エピソードが含まれている。
これらのアノテーションに基づいて、RoboInter-VQAは空間的および時間的エンボディされたVQAタスクを導入し、RoboInter-VLMの中間表現推論能力を改善した。
RoboInter-VLAはさらに、このような表現が暗黙的で明示的でモジュール型のプラン-then-executeパラダイムによるアクション実行の恩恵について研究している。
物理世界をより良くモデル化するために,中間表現を構造化条件信号として活用して将来の世界状態の制御可能な予測を行うRoboInter-Worldを紹介する。
大規模な評価は、RoboInter1.5が中間表現のための一貫した時空間足場を提供することを示している。
中間表現を単に解釈可能な信号として扱うのではなく、RoboInter1.5はそれらを双方向インターフェースとして概念化し、低レベルなアクション空間を正規化し、オープンワールド物理シミュレータの潜在ロールアウトを制限する。
関連論文リスト
- RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation [104.68774434699158]
RoboInter Manipulation Suiteはデータ、ベンチマーク、中間表現のモデルを含む統一されたリソースである。
多様な表現の半自動アノテーションを可能にする軽量GUIであるRoboInter-Toolと、571の多様なシーンにわたる230万回以上のエピソードを含む大規模なデータセットであるRoboInter-Dataで構成されている。
RoboInter-VLAは、モジュールとエンドツーエンドのVLAバリアントをサポートする、統合されたプラン-then-executeフレームワークを提供する。
論文 参考訳(メタデータ) (2026-02-10T17:01:54Z) - MobileManiBench: Simplifying Model Verification for Mobile Manipulation [70.30578259859512]
MobileManiBenchは、モバイルベースのロボット操作のための大規模なベンチマークである。
MobileManiBenchには、2つのモバイルプラットフォーム(パラレルグリッパーとデキソラスハンドロボット)、2つの同期カメラ(頭と右手首)、630のオブジェクト(オープン、クローズ、プル、プッシュ、ピック)、5つのスキル(オープン、クローズ、プッシュ、ピック)、100以上のタスクが現実的なシーンで実行される。
論文 参考訳(メタデータ) (2026-02-05T02:49:52Z) - RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics [67.11221574129937]
空間参照は、3D物理世界と相互作用するエンボディロボットの基本的な能力である。
本稿では,まず空間的理解を正確に行うことのできる3次元VLMであるRoboReferを提案する。
RoboReferは、強化微調整による一般化された多段階空間推論を推進している。
論文 参考訳(メタデータ) (2025-06-04T17:59:27Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。