論文の概要: Demonstrating Generalization Failures via Mixtures of Conditional Policies
- arxiv url: http://arxiv.org/abs/2607.03478v1
- Date: Fri, 03 Jul 2026 16:43:59 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-07 22:26:29.629377
- Title: Demonstrating Generalization Failures via Mixtures of Conditional Policies
- Title(参考訳): 条件付き政策の混合による一般化失敗の実証
- Authors: Jou Barzdukas, Jack Peck, Julian Schulz, Paulius Rauba, Steven Basart, Lennie Wells,
- Abstract要約: 強化学習(Reinforcement Learning, RL)で学習すると、制御可能な方法で一般化できない言語モデルを構築するための、シンプルで柔軟な方法を提案する。
RL トレーニングは,トレーニング分布において最高の報酬を得るポリシーを選択する。
我々はまた、将来の言語モデルで一般化が失敗する可能性のある2つの新しい方法を説明するために、我々の構成を用いる。
- 参考スコア(独自算出の注目度): 5.640047328020258
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Post-training of frontier language models is conducted on curated task suites, and inevitably leaves a distribution shift between training and deployment environments. This exposes developers to generalization failures, which are relatively poorly understood. To better understand such generalization failures, we believe the community should construct clean demonstrations under simplified conditions. To facilitate this, we propose a simple and flexible way to construct language models which fail to generalize in controllable ways when subsequently trained with Reinforcement Learning (RL) on a given distribution of training tasks. Our construction uses Supervised Fine-Tuning on a dataset of a mixture of transcripts corresponding to a collection of 'conditional policies', which can each independently be assigned certain behaviors on each different task distribution, to obtain a model that is then well approximated as a 'mixture of conditional policies.' We observe that RL training then selects for policies that obtain the highest reward on the training distribution. This can produce striking behaviors: in a controlled setting with two distributions containing identical questions prepended with two different 'trigger strings', RL training on either distribution actively degrades performance on the other to zero, even though the underlying task is identical. We also use our construction to illustrate two novel ways in which generalization may fail in future language models, corresponding to distribution shifts of task coverage and temporal context respectively. While our construction is deliberately simple and may not closely resemble 'natural' generalization failures, the resulting 'model organisms' are of interest for alignment stress-testing and generalization science, and can be used as existence proofs that training success and generalization can come apart in structured ways.
- Abstract(参考訳): フロンティア言語モデルの後のトレーニングは、キュレートされたタスクスイート上で行われ、必然的に、トレーニングとデプロイメント環境の間の分散シフトを残します。
これにより、開発者は、比較的理解されていない一般化の失敗に晒される。
このような一般化の失敗をよりよく理解するために、コミュニティは単純化された条件の下でクリーンなデモを構築するべきであると信じている。
そこで本研究では,Reinforcement Learning (RL) を用いて学習すると,制御可能な方法で一般化できない言語モデルを構築するための,シンプルで柔軟な手法を提案する。
提案手法では,各タスク分布に個別に特定の動作を割り当てることのできる「条件ポリシーの混合」とよく近似されたモデルを得るために,「条件ポリシー」の集合に対応する書き起こしのデータセットにSupervised Fine-Tuningを用いている。
RL トレーニングは,トレーニング分布において最高の報酬を得るポリシーを選択する。
2つの異なる'トリガー文字列'で予測された同一の質問を含む2つの分布を含む制御された環境では、基礎となるタスクが同一であっても、どちらかの分布に関するRLトレーニングは、もう一方のパフォーマンスをゼロに積極的に低下させる。
また,タスクカバレッジと時間的コンテキストの分布変化に対応して,将来の言語モデルで一般化が失敗する可能性のある2つの新しい方法を示す。
我々の構造は意図的に単純であり、「自然」の一般化失敗とあまり似ていないかもしれないが、結果として生じる「モデル生物」は、アライメントストレステストと一般化科学に関心を持ち、成功と一般化が構造化された方法で分解されることを示す存在証明として利用することができる。
関連論文リスト
- Robust Finetuning of Vision-Language-Action Robot Policies via Parameter Merging [53.41119829581115]
大規模で多様なデータセットに基づいて訓練された汎用ロボットポリシーは、一般化する能力を実証している。
トレーニングデータに含まれていない新しいタスクにはまだ不足しています。
本研究では,ファインタニング時の一般政策の一般化能力を保全する手法を開発した。
論文 参考訳(メタデータ) (2025-12-09T08:02:11Z) - Self-Predictive Representations for Combinatorial Generalization in Behavioral Cloning [20.1753113722028]
後続表現による長距離時間一貫性がいかに一般化を促進するかを示す。
次に,GCBCのための単純な表現学習目的である$textBYOL-gamma$を提案する。
論文 参考訳(メタデータ) (2025-06-11T19:32:41Z) - Any-Shift Prompting for Generalization over Distributions [66.29237565901734]
即時学習におけるトレーニングとテスト分布の関係を考察する一般的な確率的推論フレームワークである「任意のシフトプロンプト」を提案する。
このフレームワーク内では、テストプロンプトが分散関係を利用して、CLIPイメージ言語モデルのトレーニングからテストディストリビューションへの一般化を導く。
ネットワークは、トレーニング情報とテスト情報の両方をフィードフォワードパスに組み込んだ調整されたテストプロンプトを生成し、テスト時の追加のトレーニングコストを回避する。
論文 参考訳(メタデータ) (2024-02-15T16:53:42Z) - Class Distribution Shifts in Zero-Shot Learning: Learning Robust Representations [3.8980564330208662]
本研究では,前もって変化の原因となる属性が不明であると仮定したモデルを提案し,解析する。
提案アルゴリズムは,実世界のデータセット上でのシミュレーションと実験の両方において,多様なクラス分布への一般化を改善する。
論文 参考訳(メタデータ) (2023-11-30T14:14:31Z) - Time-series Generation by Contrastive Imitation [87.51882102248395]
モーメントマッチングの目的によってモチベーションされ、複合的エラーを軽減し、局所的(しかし前方的な)遷移ポリシーを最適化する。
推論において、学習されたポリシーは反復的なサンプリングのジェネレータとして機能し、学習されたエネルギーはサンプルの品質を評価するための軌道レベル尺度として機能する。
論文 参考訳(メタデータ) (2023-11-02T16:45:25Z) - Self-regulating Prompts: Foundational Model Adaptation without
Forgetting [112.66832145320434]
本稿では,PromptSRCと呼ばれる自己正規化フレームワークを提案する。
PromptSRCはタスク固有の汎用表現とタスクに依存しない汎用表現の両方に最適化するプロンプトを導く。
論文 参考訳(メタデータ) (2023-07-13T17:59:35Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。