論文の概要: Exploration and Online Transfer with Behavioral Foundation Models
- arxiv url: http://arxiv.org/abs/2606.29980v2
- Date: Tue, 30 Jun 2026 08:37:47 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-01 13:50:27.731495
- Title: Exploration and Online Transfer with Behavioral Foundation Models
- Title(参考訳): 行動基礎モデルによる探索とオンライン移動
- Abstract要約: 我々は,このオンライン学習問題を,盗賊的探索探索問題の観点から考えることが可能であることを示す。
上信頼境界から着想を得た定式化を導出し、不確実行列の固有値の最小化により探索が可能であることを示す。
- 参考スコア(独自算出の注目度): 3.015354201629835
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Zero-shot Transfer in Reinforcement Learning (RL) aims to train an agent that can generate optimal policies for any reward function, without additional learning at transfer time, while training only on reward-free trajectories. For their generality over tasks, such models are sometimes called ``Behavioral Foundation Models'' (BFMs). While they have shown strong performances and improvements in recent years, the current framework and algorithms still assume that, during the transfer phase, the agent is informed offline about the reward (the task to solve) through a dataset of state-reward pairs, which it uses to pick the best policy to deploy. However, in practice if the reward is a black-box (e.g. direct user feedback), it is not possible to generate such a dataset: it is necessary to observe the reward through interactions with the environment. In other words, the current framework of offline transfer is not aligned with the traditional RL setting of online learning through trial-and-error, which requires exploration in order to find rewards. This paper proposes to tackle this new online transfer in zero-shot RL, with the key insight that the BFM itself can be used to generate exploration policies. We show that it is possible to frame this online learning problem in terms of a bandit-like exploration-exploitation problem. More precisely, at each step the bandit algorithm recommends a policy, the BFM executes it in the environment, which yields a reward and a new state; we repeat the process until we converge to the optimal policy. In the popular context of linear reward approximation, we derive a formulation inspired by Upper Confidence Bound and show that exploration can be achieved through the minimization of the eigenvalues of an uncertainty matrix. We evaluate qualitatively and quantitatively our framework on a simple environment to validate the concept of our method.
- Abstract(参考訳): 強化学習におけるゼロショット伝達(Zero-shot Transfer in Reinforcement Learning, RL)は、報酬のない軌道のみを訓練しながら、転送時に追加学習することなく、報酬関数に対して最適なポリシーを生成できるエージェントを訓練することを目的としている。
タスクに対する一般性については、そのようなモデルを `Behavioral Foundation Models' (BFMs) と呼ぶこともある。
彼らは近年、強力なパフォーマンスと改善を示しているが、現在のフレームワークとアルゴリズムは、移行フェーズの間、エージェントは、デプロイする最良のポリシーを選択するために使用するステート-リワードペアのデータセットを通じて、報酬(解決すべきタスク)についてオフラインで通知されると仮定している。
しかし、実際には報酬がブラックボックス(例えばユーザからの直接フィードバック)であれば、そのようなデータセットを生成することはできない。
言い換えれば、現在のオフライン転送のフレームワークは、試行錯誤によるオンライン学習の従来のRL設定と一致していない。
本稿では,この新たなオンライン転送をゼロショットRLで行うことを提案する。
我々は,このオンライン学習問題を,盗賊的探索探索問題の観点から考えることが可能であることを示す。
より正確には、バンディットアルゴリズムはポリシーを推奨し、BFMはそれを環境内で実行し、報酬と新しい状態を与えます。
線形報酬近似の一般的な文脈では、アッパー信頼境界に着想を得た定式化を導出し、不確実行列の固有値の最小化により探索が可能であることを示す。
簡単な環境下での質的,定量的なフレームワークの評価を行い,提案手法の概念を検証した。
関連論文リスト
- Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation [50.696688705287755]
我々は、強化学習におけるスパース報酬課題を克服するために、相互情報自己評価を提案する。
MISEにより、エージェントは、疎外的信号を補う高密度な内部報酬から自律的に学習することができる。
我々は、後見自己評価報酬を利用することは、政策と代行報酬政策の間のKL分散項と相互情報を組み合わせた目的を最小化することと等価であることを示す。
論文 参考訳(メタデータ) (2026-04-13T15:18:51Z) - Maximum Entropy Behavior Exploration for Sim2Real Zero-Shot Reinforcement Learning [18.905272125661824]
実際のロボットシステムにおける四足歩行制御のために,$textitonline$zero-shot RLについて検討した。
我々は、教師なしの行動探索戦略と正規化評論家を組み合わせたオンラインゼロショットRLアルゴリズムであるFB-MEBEを紹介する。
論文 参考訳(メタデータ) (2026-03-26T14:07:01Z) - Agentic Reinforcement Learning with Implicit Step Rewards [92.26560379363492]
大規模言語モデル (LLMs) は強化学習 (agentic RL) を用いた自律的エージェントとして発展している。
我々は,標準RLアルゴリズムとシームレスに統合された一般的なクレジット割り当て戦略であるエージェントRL(iStar)について,暗黙的なステップ報酬を導入する。
我々は,WebShopとVisualSokobanを含む3つのエージェントベンチマークと,SOTOPIAにおける検証不可能な報酬とのオープンなソーシャルインタラクションについて評価した。
論文 参考訳(メタデータ) (2025-09-23T16:15:42Z) - TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning [48.31236495564408]
本稿では,TROFI(Trjectory-Ranked Offline Inverse reinforcement Learning)を提案する。
TROFIは、事前に定義された報酬関数なしでオフラインでポリシーを効果的に学習するための新しいアプローチである。
TROFIは基準線を一貫して上回り、基本真理報酬を用いてポリシーを学ぶのに相容れない性能を示す。
論文 参考訳(メタデータ) (2025-06-27T08:22:41Z) - Best Policy Learning from Trajectory Preference Feedback [11.896067099790962]
推論ベースの強化学習(PbRL)は、より堅牢な代替手段を提供する。
本稿では, PbRLにおける最適政策識別問題について検討し, 生成モデルの学習後最適化を動機とした。
本稿では,Top-Two Thompson Smplingにヒントを得た新しいアルゴリズムであるPosterior Smpling for Preference Learning(mathsfPSPL$)を提案する。
論文 参考訳(メタデータ) (2025-01-31T03:55:10Z) - MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization [91.80034860399677]
強化学習アルゴリズムは、現在のベスト戦略の活用と、より高い報酬につながる可能性のある新しいオプションの探索のバランスを図ることを目的としている。
我々は本質的な探索と外生的な探索のバランスをとるためのフレームワークMaxInfoRLを紹介する。
提案手法は,マルチアームバンディットの簡易な設定において,サブリニアな後悔を実現するものである。
論文 参考訳(メタデータ) (2024-12-16T18:59:53Z) - Reinforcement Learning from Bagged Reward [46.16904382582698]
強化学習(RL)では、エージェントが取るアクション毎に即時報奨信号が生成されることが一般的である。
多くの実世界のシナリオでは、即時報酬信号の設計は困難である。
本稿では,双方向の注意機構を備えた新たな報酬再分配手法を提案する。
論文 参考訳(メタデータ) (2024-02-06T07:26:44Z) - Model-based Safe Deep Reinforcement Learning via a Constrained Proximal
Policy Optimization Algorithm [4.128216503196621]
オンライン方式で環境の遷移動態を学習する,オンライン型モデルに基づくセーフディープRLアルゴリズムを提案する。
我々は,本アルゴリズムがより標本効率が高く,制約付きモデルフリーアプローチと比較して累積的ハザード違反が低いことを示す。
論文 参考訳(メタデータ) (2022-10-14T06:53:02Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。