論文の概要: RL Bootstrapping of OpenVLA-OFT for a Novel Robot Embodiment
- arxiv url: http://arxiv.org/abs/2608.01013v1
- Date: Sun, 02 Aug 2026 05:33:36 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-04 15:07:25.053048
- Title: RL Bootstrapping of OpenVLA-OFT for a Novel Robot Embodiment
- Title(参考訳): OpenVLA-OFTのRLブートストラップによる新しいロボット・エボディメント
- Authors: Damir Nurtdinov, Alexei Kornaev, Alexander Maloletov,
- Abstract要約: ケーブル駆動並列ロボット上でのOpenVLA-OFTのゼロデモエボディメントアライメントアライメント(ゼロデモエボディメントアライメントアライメントアライメント)について検討する。
シュミレーション状態から計算した濃密な幾何報酬を用いたシミュレーションにおいて強化学習を用いる。
結果はまだ堅牢な操作を確立していないが、RLのみのブートストラッピングが最初の使用可能な言語条件のコントローラを作成できるという強い証拠を提供する。
- 参考スコア(独自算出の注目度): 41.99844472131922
- License: http://creativecommons.org/publicdomain/zero/1.0/
- Abstract: Adapting a pretrained vision-language-action (VLA) policy to a new robot usually assumes embodiment-specific demonstrations. This assumption is especially restrictive for custom robots whose morphology differs strongly from the manipulators seen in large robot datasets. We study a harder setting: zero-demo embodiment alignment of OpenVLA-OFT on a cable-driven parallel robot (CDPR) with a simple gripper and a previously unseen control interface. Instead of supervised fine-tuning, we use reinforcement learning in simulation with dense geometric rewards computed from simulator state. The training is performed in two stages: a PPO stage for directional motion primitives, followed by GRPO continuation from the PPO checkpoint with an expanded instruction space that includes object-conditioned commands. On the four shared directional instructions, the average held-out success rate improves from 34.25\% after PPO to 53.50\% after PPO$\rightarrow$GRPO, with especially large gains on \texttt{move left} and \texttt{move backward}. In the GRPO stage we additionally introduce \texttt{move to <object>} over eight target objects and obtain 39/400 = 9.75\% strict success, while qualitative rollouts frequently show correct target-directed approach behavior before late-stage instability. Compared with prior OpenVLA and OpenVLA-OFT results, which rely on demonstration datasets and mostly standard rigid-arm embodiments, our method uses no embodiment-specific dataset at all. The results do not yet establish robust manipulation, but they provide stronger evidence that RL-only bootstrapping can create the first usable language-conditioned controller for a genuinely novel embodiment.
- Abstract(参考訳): 事前訓練された視覚言語アクション(VLA)ポリシーを新しいロボットに適用することは、通常、エンボディメント固有のデモンストレーションを前提とします。
この仮定は、大きなロボットデータセットに見られるマニピュレータと形態が強く異なるカスタムロボットに対して特に制限的である。
ケーブル駆動並列ロボット(CDPR)におけるOpenVLA-OFTのゼロデモエボディメントアライメントアライメント(ゼロデモエボディメントアライメントアライメントアライメントアライメント)について検討した。
教師付き微調整の代わりに、シミュレータの状態から計算された密度の幾何報酬を用いたシミュレーションに強化学習を用いる。
トレーニングは、方向運動プリミティブのためのPPOステージと、オブジェクト条件のコマンドを含む拡張命令空間を備えたPPOチェックポイントからのGRPO継続の2段階で行われる。
4つの共有方向指示では、平均ホールトアウト成功率は、PPO後の34.25\%から、PPO$\rightarrow$GRPO後の53.50\%に改善され、特に \texttt{move left} と \texttt{move backward} では大きな利得となる。
GRPO の段階では,8つの対象対象に対して \texttt{move to <object>} を導入し,39/400 = 9.75\% の厳密な成功を得た。
従来のOpenVLAとOpenVLA-OFTの結果は、デモデータセットや標準の剛腕エボディメントに依存しているが、本手法ではエボディメント固有のデータセットは一切使用していない。
結果はまだ堅牢な操作を確立していないが、RLのみのブートストラッピングが真に新しい実施のための最初の使用可能な言語コンディショナブルコントローラを作成できるという強い証拠を提供する。
関連論文リスト
- AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models [9.608633915316252]
VLA(Vision-Language-Action)モデルでは、一般化可能なロボット操作の可能性を示している。
現在のパラダイムは、教師付き微調整中の粗大でハイレベルなタスク命令に依存している。
スケーラブルなオフライン後トレーニングパイプラインと統合された,最初のサブタスク対応VLAフレームワークである方法を提案する。
論文 参考訳(メタデータ) (2026-03-09T15:52:48Z) - Universal Pose Pretraining for Generalizable Vision-Language-Action Policies [83.39008378156647]
既存のVision-Language-Action(VLA)モデルは、しばしば機能崩壊と訓練効率の低下に悩まされる。
本稿では,VLAトレーニングを3次元空間前駆体抽出のための事前学習フェーズに分離する,分離されたパラダイムであるPose-VLAを提案する。
我々のフレームワークは2段階の事前学習パイプラインに従い、ポーズと動きのアライメントによる基本的な空間接地を確立する。
論文 参考訳(メタデータ) (2026-02-23T11:00:08Z) - TGRPO :Fine-tuning Vision-Language-Action Model via Trajectory-wise Group Relative Policy Optimization [12.061547251822326]
Trajectory-based Group Relative Policy Optimization (TGRPO)は、Visual-Language-Action(VLA)モデルのためのオンラインRLベースのトレーニングフレームワークである。
TGRPOの平均成功率は80.7%で、これはスーパーバイザードファインチューニング(SFT)よりも4.2%高く、他の代表的RLベースのポストトレーニング手法よりも優れていた。
論文 参考訳(メタデータ) (2025-06-10T04:27:49Z) - HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation [54.03004125910057]
階層型視覚-言語-アクションモデルは、標準的なモノリシックVLAモデルよりも、ドメイン外のデータを利用するのに効果的であることを示す。
階層設計により、高レベルなVLMは、オフドメイン微調整データと実ロボットテストシナリオの間の重要なドメインギャップをまたいで転送可能であることを示す。
論文 参考訳(メタデータ) (2025-02-08T07:50:22Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。