Fugu-MT 論文翻訳(概要): Learning while Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies

論文の概要: Learning while Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies

arxiv url: http://arxiv.org/abs/2605.00416v1
Date: Fri, 01 May 2026 05:20:26 GMT
ステータス: 翻訳完了
システム内更新日: 2026-05-04 17:43:28.856693
Title: Learning while Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies
Title（参考訳）: デプロイ中の学習:ジェネラリストロボット政策のためのフリートスケール強化学習
Authors: Yi Wang, Xinchen Li, Pengwei Xie, Pu Yang, Buqing Nie, Yunuo Cai, Qinglin Zhang, Chendi Qu, Jeffrey Wu, Jianheng Song, Xinlin Ren, Jingshun Huang, Mingjie Pan, Siyuan Feng, Zhi Chen, Jianlan Luo,
Abstract要約: 汎用的なロボットポリシーは、大規模な事前トレーニングの恩恵を受ける傾向にあるが、オフラインデータだけでは、堅牢な現実世界のデプロイメントには不十分である。本稿では,VLA(Vision-Language-Action)ポリシーの継続学習のための,艦隊規模のオフライン-オンライン強化学習フレームワークであるLWDを紹介する。
参考スコア（独自算出の注目度）: 23.266003019334438
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Abstract: Generalist robot policies increasingly benefit from large-scale pretraining, but offline data alone is insufficient for robust real-world deployment. Deployed robots encounter distribution shifts, long-tail failures, task variations, and human correction opportunities that fixed demonstration datasets cannot fully capture. We present Learning While Deploying (LWD), a fleet-scale offline-to-online reinforcement learning framework for continual post-training of generalist Vision-Language-Action (VLA) policies. Starting from a pretrained VLA policy, LWD closes the loop between deployment, shared physical experience, policy improvement, and redeployment by using autonomous rollouts and human interventions collected across a robot fleet. To stabilize learning from heterogeneous, sparse-reward fleet data, LWD combines Distributional Implicit Value Learning (DIVL) for robust value estimation with Q-learning via Adjoint Matching (QAM) for policy extraction in flow-based VLA action generators. We validate LWD on a fleet of 16 dual-arm robots across eight real-world manipulation tasks, including semantic grocery restocking and 3--5 minute long-horizon tasks. A single generalist policy improves as fleet experience accumulates, reaching an average success rate of 95%, with the largest gains on long-horizon tasks.
Abstract（参考訳）: 汎用的なロボットポリシーは、大規模な事前トレーニングの恩恵を受ける傾向にあるが、オフラインデータだけでは、堅牢な現実世界のデプロイメントには不十分である。デプロイされたロボットは、分散シフト、ロングテール障害、タスクのバリエーション、固定されたデモデータセットが完全にキャプチャできない人間の修正機会に遭遇する。本稿では,VLA(Vision-Language-Action)ポリシーの継続学習のための,艦隊規模のオフライン-オンライン強化学習フレームワークであるLWDを紹介する。事前訓練されたVLAポリシから始めて、LWDは、ロボット群全体で収集された自律的なロールアウトと人間の介入を使用することで、デプロイメント、共有された物理的エクスペリエンス、ポリシー改善、再デプロイの間のループを閉じる。不均一でスパース・リワードな艦隊データからの学習を安定させるために、LWDは分散インプリシット・バリュー・ラーニング(DIVL)と、フローベースのVLAアクションジェネレータのポリシー抽出のための随伴マッチング(QAM)によるQ-ラーニングを組み合わせた。 LWDは、セマンティック・グロサリー・リストックや3～5分のロングホライゾン・タスクを含む8つの実世界の操作タスクにまたがる16のデュアルアーム・ロボット群で検証する。単一のジェネラリスト政策は、艦隊経験が蓄積するにつれて改善され、平均95%の成功率に達する。

論文の概要: Learning while Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies

関連論文リスト