論文の概要: PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations
- arxiv url: http://arxiv.org/abs/2604.27472v1
- Date: Thu, 30 Apr 2026 06:14:02 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-05-01 16:31:53.951522
- Title: PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations
- Title(参考訳): PRTS:コントラスト表現による原始推論とタスクシステム
- Authors: Yang Zhang, Jiangyuan Zhao, Chenyou Fan, Fangzheng Yan, Tian Li, Haitong Tang, Sen Fu, Xuan'er Wu, Qizhen Weng, Weinan Zhang, Xiu Li, Chi Zhang, Chenjia Bai, Xuelong Li,
- Abstract要約: 我々は,目標達成型強化学習を通じて事前学習を再構築するVLA基盤モデルであるtextbfPRTS(textbfPrimitive textbfReasoning and textbfTasking textbfSystem)を提案する。
- 参考スコア(独自算出の注目度): 66.94988600664574
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Vision-Language-Action (VLA) models advance robotic control via strong visual-linguistic priors. However, existing VLAs predominantly frame pretraining as supervised behavior cloning, overlooking the fundamental nature of robot learning as a goal-reaching process that requires understanding temporal task progress. We present \textbf{PRTS} (\textbf{P}rimitive \textbf{R}easoning and \textbf{T}asking \textbf{S}ystem), a VLA foundation model that reformulates pretraining through Goal-Conditioned Reinforcement Learning. By treating language instructions as goals and employing contrastive reinforcement learning, PRTS learns a unified embedding space where the inner product of state-action and goal embeddings approximates the log-discounted goal occupancy, the probability of reaching the language-specified goal from the current state-action, quantitatively assessing physical feasibility beyond static semantic matching. PRTS draws this dense goal-reachability supervision directly from offline trajectories without reward annotations, and folds it into the VLM backbone via a role-aware causal mask, incurring negligible overhead over vanilla behavior cloning. This paradigm endows the high-level reasoning system with intrinsic goal reachability awareness, bridging semantic reasoning and temporal task progress, and further benefits goal-conditioned action prediction. Pretrained on 167B tokens of diverse manipulation and embodied-reasoning data, PRTS reaches state-of-the-art performance on LIBERO, LIBERO-Pro, LIBERO-Plus, SimplerEnv, and a real-world suite of 14 complex tasks, with particularly substantial gains on long-horizon, contact-rich, and zero-shot novel-instruction settings, confirming that injecting goal-reachability awareness significantly improves both execution success and long-horizon planning of general-purpose robotic foundation policies.
- Abstract(参考訳): VLA(Vision-Language-Action)モデルは、強力な視覚言語によるロボット制御を推進している。
しかしながら、既存のVLAは、時間的タスクの進捗を理解することを必要とする目標獲得プロセスとして、ロボット学習の基本的な性質を見越して、監督行動のクローンとして事前訓練を主に実施している。
本稿では、ゴール・コンディション強化学習を通じて事前学習を再構築する VLA 基礎モデルである \textbf{PRTS} (\textbf{P}rimitive \textbf{R}easoning and \textbf{T}asking \textbf{S}ystem) を提案する。
PRTSは、言語命令を目標として扱い、対照的な強化学習を採用することにより、状態アクションと目標埋め込みの内積が対数分散目標占有率に近似する統合埋め込み空間を学習し、言語特定目標に達する確率を現在の状態アクションから推定し、静的意味マッチング以上の物理的実現可能性を評価する。
PRTSは、報酬アノテーションなしでオフラインの軌跡から直接、この密集した目標到達可能性の監視を引き出し、ロール認識因果マスクを介してVLMバックボーンに折り畳み、バニラの動作クローンに対する無視できないオーバーヘッドを引き起こす。
このパラダイムは、本質的な目標到達可能性認識、意味的推論と時間的タスクの進捗を橋渡しし、さらにゴール条件付き行動予測の恩恵を与える高レベル推論システムを提供する。
167Bトークンの多彩な操作と実施データに基づいて、PRTSはLIBERO、LIBERO-Pro、LIBERO-Plus、SimplerEnv、および14の複雑なタスクからなる実世界のスイートに到達し、特にロングホライゾン、コンタクトリッチ、ゼロショットのノベル・インストラクション設定において顕著な利益を上げ、目標到達可能性の認識を注入することで、汎用ロボット基盤ポリシーの実行成功とロングホライゾン計画の両方を大幅に改善することを確認した。
関連論文リスト
- See, Plan, Rewind: Progress-Aware Vision-Language-Action Models for Robust Robotic Manipulation [59.07792608884117]
本稿では,See, Plan, Rewind (SPR)について紹介する。
SPRは、現在の状態と今後のマイルストーンを見て、次の2Dウェイポイントに向けて軌道を計画し、障害時に回復可能な状態に戻すという、継続的なコアサイクルを通じて運用される。
SPRは、OpenVLA-OFTとUniVLAを上回る最小のパフォーマンス低下で最先端のロバスト性を達成する。
論文 参考訳(メタデータ) (2026-03-10T07:22:51Z) - Self-Correcting VLA: Online Action Refinement via Sparse World Imagination [55.982504915794514]
本稿では, 自己補正VLA (SC-VLA) を提案する。
SC-VLAは最先端のパフォーマンスを達成し、最高タスクスループットを16%削減し、最高パフォーマンスのベースラインよりも9%高い成功率を得る。
論文 参考訳(メタデータ) (2026-02-25T06:58:06Z) - PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation [27.791908160098625]
PALMは、インタラクション中心のアベイランス推論とサブタスクプログレスキューに関するポリシー学習を構築する。
Palmはシミュレーションや実世界の実験において、一貫してベースラインを上回っている。
論文 参考訳(メタデータ) (2026-01-11T21:00:58Z) - Learning Affordances at Inference-Time for Vision-Language-Action Models [50.93181349331096]
ロボット工学において、VLA(Vision-Language-Action Model)は複雑な制御タスクを解くための有望な道を提供する。
本稿では,VLAの低レベルポリシーを過去の経験を条件とした高レベルVLMに接続するLITEN(Learning from Inference-Time Execution)を紹介する。
提案手法は,低レベルVLAの計画の生成と実行を行う推論フェーズと,その結果を反映した評価フェーズとを反復する。
論文 参考訳(メタデータ) (2025-10-22T16:43:29Z) - Executable Analytic Concepts as the Missing Link Between VLM Insight and Precise Manipulation [70.8381970762877]
VLM(Vision-Language Models)は、セマンティック推論とタスク計画において顕著な能力を示す。
本稿では,VLMに基づく推論を実行可能な解析概念を通じて基礎づける新しいフレームワークであるGRACEを紹介する。
G GRACEは高レベル命令理解と低レベルロボット制御の統一的で解釈可能なインターフェースを提供する。
論文 参考訳(メタデータ) (2025-10-09T09:08:33Z) - IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction [51.130510883952546]
Vision-Language-Action(VLA)モデルは、事前訓練された視覚言語モデル(VLM)を活用して、ロボット制御との認識を両立させる。
カリキュラム学習パラダイムと効率的な推論機構を備えたVLAフレームワークである textbfIntentionVLA を提案する。
提案手法はまず,意図推論,空間的接地,コンパクトな具体的推論を組み合わせ,慎重に設計した推論データを活用する。
論文 参考訳(メタデータ) (2025-10-09T04:49:46Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。