論文の概要: Instruction Quality Matters: Refining Instructions for Effective Preference Learning
- arxiv url: http://arxiv.org/abs/2608.26779v1
- Date: Thu, 27 Aug 2026 08:09:36 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-28 16:30:58.302025
- Title: Instruction Quality Matters: Refining Instructions for Effective Preference Learning
- Title(参考訳): 授業品質の諸問題:効果的な選好学習のための指導を精査する
- Authors: Seohyeong Lee, Hwaran Lee, Buru Chang,
- Abstract要約: 選好学習において,指導の質を隠されたボトルネックとして認識する。
提案手法は, サンプル応答品質の天井面と床面の両方に制約があることを示す。
本稿では、報酬信号を用いて弱い命令を選択して修正する命令補充パイプラインを提案する。
- 参考スコア(独自算出の注目度): 19.674135177469186
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Preference learning optimizes models using response pairs, yet the informativeness of these pairs is fundamentally shaped by the instructions from which they are generated. We identify instruction quality as a hidden bottleneck in preference learning: low-quality or ambiguous instructions restrict the response-quality distribution, limiting strong chosen responses and weakening preference signals. Through Best- and Worst-of-N analyses, we show that instruction quality constrains both the ceiling and floor of sampled response quality. Motivated by this observation, we introduce an instruction-refinement pipeline that selects weak instructions using reward signals and revises them with rubric-guided LLM feedback, improving preference data without discarding examples. Across offline and online preference learning settings, experiments on multiple models and benchmarks show broad alignment improvements over original data and alternative data-improvement strategies. Further analyses indicate that instruction refinement raises achievable response quality and complements response-centric preference data curation. Overall, instruction quality emerges as a key factor governing how informative preference signals are formed for LLM alignment. Code is available at: https://github.com/01choco/instruction-refinement/
- Abstract(参考訳): 優先順位学習は、応答ペアを使用してモデルを最適化するが、これらのペアの情報性は、それらが生成される命令によって根本的に形作られる。
低品質または曖昧な指示は、応答品質の分布を制限し、強い選択された応答を制限し、選好シグナルを弱める。
ベスト・オブ・N分析により, サンプル応答品質の天井と床の両方に指示品質が制約されることが示唆された。
この観測により,報奨信号を用いて弱い命令を選択して,豪華なLLMフィードバックで修正し,例を捨てることなく好みデータを改善する命令補充パイプラインが導入された。
オフラインおよびオンラインの嗜好学習設定全体において、複数のモデルとベンチマークの実験は、オリジナルのデータと代替データ改善戦略よりも広範なアライメント改善を示している。
さらに分析したところ、命令の洗練は達成可能な応答品質を高め、応答中心の嗜好データキュレーションを補完することが示された。
全体として、LLMアライメントのための情報優先信号の作り方を決定する重要な要因として、命令品質が出現する。
コードは、https://github.com/01choco/instruction-refinement/で入手できる。
関連論文リスト
- Aligning Large Language Models to Follow Instructions and Hallucinate Less via Effective Data Filtering [66.5524727179286]
NOVAは、幻覚を減らすための学習知識とよく一致した高品質なデータを特定するために設計されたフレームワークである。
内部整合性探索(ICP)とセマンティック等価同定(SEI)が含まれており、LLMが命令データとどれだけ親しみやすいかを測定する。
選択したサンプルの品質を確保するため,親しみ以上の特性を考慮した専門家による報酬モデルを導入する。
論文 参考訳(メタデータ) (2025-02-11T08:05:56Z) - Reward-Augmented Data Enhances Direct Preference Alignment of LLMs [63.32585910975191]
報奨条件付き大言語モデル(LLM)を導入し、データセット内の応答品質のスペクトル全体から学習する。
当社のアプローチは,DPOをかなりのマージンで継続的に向上させることを示す。
本手法は,嗜好データの有用性を最大化するだけでなく,未学習の問題も軽減し,データ拡張を超えてその広範な効果を実証する。
論文 参考訳(メタデータ) (2024-10-10T16:01:51Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。