論文の概要: RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation
- arxiv url: http://arxiv.org/abs/2607.18709v2
- Date: Wed, 22 Jul 2026 05:33:17 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-23 14:36:03.346134
- Title: RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation
- Title(参考訳): RoboInter1.5: 身体的世界モデリングとロボットマニピュレーションのためのホロスティックな中間表現スイート
- Abstract要約: 本稿では,ロボット操作と擬似世界モデリングの中間表現を拡張し,包括的に表現するRoboInter1.5について紹介する。
RoboInter1.5は、厳密な操作指向の中間表現を中心としたデータ、ベンチマーク、モデルの統一されたリソースを提供する。
- 参考スコア(独自算出の注目度): 46.68319840674789
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Existing robot datasets remain expensive to curate, embodiment-specific, and insufficiently annotated with the fine-grained structure required for generalizable reasoning, execution, or long-horizon environment dynamics simulation. Building on our prior work, RoboInter1.0, we present RoboInter1.5, an extended and holistic suite of intermediate representations for both robotic manipulation and embodied world modeling. RoboInter1.5 provides a unified resource of data, benchmarks, and models centered on dense manipulation-oriented intermediate representations. Specifically, RoboInter-Data contains over 230k manipulation episodes across 571 scenes with dense per-frame annotations covering more than ten types of intermediate representations, including subtasks, primitive skills, object and gripper grounding, segmentation, affordance, grasp poses, contact points, motion traces, etc. Built upon these annotations, RoboInter-VQA introduces spatial and temporal embodied VQA tasks to benchmark and improve the intermediate-representation reasoning capabilities of our RoboInter-VLM. RoboInter-VLA further studies how such representations benefit action execution through implicit, explicit, and modular plan-then-execute paradigms. To better model the physical world, we further introduce RoboInter-World, which leverages intermediate representations as structured conditioning signals for controllable prediction of future world states. Extensive evaluations demonstrate that RoboInter1.5 provides a unified spatiotemporal scaffolding for intermediate representations. Rather than treating intermediate representations merely as interpretable signals, RoboInter1.5 conceptualizes them as a bidirectional interface that both regularizes low-level action spaces and constrains the latent rollouts of open-world physical simulators.
- Abstract(参考訳): 既存のロボットデータセットは、一般的な推論、実行、長期の環境力学シミュレーションに必要なきめ細かい構造で、キュレートし、具体化し、十分に注釈を付けるのに高価である。
ロボット操作と擬似世界モデリングの両方のための中間表現を拡張し包括的に表現するRoboInter1.5を,我々の以前の研究であるRoboInter1.0に基づいて紹介する。
RoboInter1.5は、厳密な操作指向の中間表現を中心としたデータ、ベンチマーク、モデルの統一されたリソースを提供する。
具体的には、RoboInter-Dataには、サブタスク、プリミティブスキル、オブジェクトとグリップグラウンド、セグメンテーション、セグメンテーション、グリップポーズ、コンタクトポイント、モーショントレースなどを含む10種類以上の中間表現を含む、フレーム単位のアノテーションを備えた571シーンにわたる230kの操作エピソードが含まれている。
これらのアノテーションに基づいて、RoboInter-VQAは空間的および時間的エンボディされたVQAタスクを導入し、RoboInter-VLMの中間表現推論能力を改善した。
RoboInter-VLAはさらに、このような表現が暗黙的で明示的でモジュール型のプラン-then-executeパラダイムによるアクション実行の恩恵について研究している。
物理世界をより良くモデル化するために,中間表現を構造化条件信号として活用して将来の世界状態の制御可能な予測を行うRoboInter-Worldを紹介する。
大規模な評価は、RoboInter1.5が中間表現のための一貫した時空間足場を提供することを示している。
中間表現を単に解釈可能な信号として扱うのではなく、RoboInter1.5はそれらを双方向インターフェースとして概念化し、低レベルなアクション空間を正規化し、オープンワールド物理シミュレータの潜在ロールアウトを制限する。
関連論文リスト
- AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding [55.535374902204296]
AffordanceVLAは、タスク指向の中間表現として構造化された価格予測を導入する統一フレームワークである。
AffordanceVLAは様々な操作シナリオで高いパフォーマンスを実現していることを示す。
論文 参考訳(メタデータ) (2026-06-04T13:28:51Z) - The Yes-Man Syndrome: Benchmarking Abstention in Embodied Robotic Agents [14.695254264082273]
視覚言語モデル(VLM)は、エンボディエージェントのプランナーとして使用される。
本稿では, ロボット工学の文脈において, 禁忌を分類するための分類法を提案する。
本稿では,画像に接地した禁忌指示を生成するためのフレームワークであるRoboAbstentionを紹介する。
論文 参考訳(メタデータ) (2026-05-19T22:32:44Z) - RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation [104.68774434699158]
RoboInter Manipulation Suiteはデータ、ベンチマーク、中間表現のモデルを含む統一されたリソースである。
多様な表現の半自動アノテーションを可能にする軽量GUIであるRoboInter-Toolと、571の多様なシーンにわたる230万回以上のエピソードを含む大規模なデータセットであるRoboInter-Dataで構成されている。
RoboInter-VLAは、モジュールとエンドツーエンドのVLAバリアントをサポートする、統合されたプラン-then-executeフレームワークを提供する。
論文 参考訳(メタデータ) (2026-02-10T17:01:54Z) - MobileManiBench: Simplifying Model Verification for Mobile Manipulation [70.30578259859512]
MobileManiBenchは、モバイルベースのロボット操作のための大規模なベンチマークである。
MobileManiBenchには、2つのモバイルプラットフォーム(パラレルグリッパーとデキソラスハンドロボット)、2つの同期カメラ(頭と右手首)、630のオブジェクト(オープン、クローズ、プル、プッシュ、ピック)、5つのスキル(オープン、クローズ、プッシュ、ピック)、100以上のタスクが現実的なシーンで実行される。
論文 参考訳(メタデータ) (2026-02-05T02:49:52Z) - InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation [77.07565723756119]
InternVLA-A1は動的予測機能を備えた視覚言語モデルである。
我々は、実世界のロボットデータ、合成シミュレーションデータ、人間のビデオなどを用いて、これらのモデルを異種データソース上で事前訓練する。
InternVLA-A1を実世界の12のロボットタスクとシミュレーションベンチマークで評価した。
論文 参考訳(メタデータ) (2026-01-05T18:54:29Z) - RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics [67.11221574129937]
空間参照は、3D物理世界と相互作用するエンボディロボットの基本的な能力である。
本稿では,まず空間的理解を正確に行うことのできる3次元VLMであるRoboReferを提案する。
RoboReferは、強化微調整による一般化された多段階空間推論を推進している。
論文 参考訳(メタデータ) (2025-06-04T17:59:27Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。