論文の概要: EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation
- arxiv url: http://arxiv.org/abs/2605.26785v1
- Date: Tue, 26 May 2026 09:54:53 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-05-27 17:51:41.863008
- Title: EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation
- Title(参考訳): EmoDistill: 対人交渉における言語モデルエージェントのオフライン感情スキル蒸留
- Authors: Yunbo Long, Haolang Zhao, Lukas Beckenbauer, Liming Xu, Alexandra Brintrup,
- Abstract要約: 言語モデルエージェントに感情交渉スキルを蒸留するオフラインフレームワークである textbfEmoDistill を導入する。
EmoDistillは感情戦略を感情選択と感情表現に分解する。
4つの感情に敏感なネゴシエーションドメインで、EmoDistillフレームワークの下でトレーニングされたSLMポリシーが最も有効である。
- 参考スコア(独自算出の注目度): 44.481959726034574
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Post-trained LLMs are often optimized to align responses with human preferences, making them safe, polite, and conversationally appropriate. In adversarial negotiation, however, this alignment can become a vulnerability: emotionally framed language may steer agents toward the counterparty's interests. Using GoEmotions-based affective prompting, we show that emotion substantially shifts negotiation outcomes, suggesting that emotion is a strategic action channel rather than a surface style. Thus, we introduce \textbf{EmoDistill}, an offline framework for distilling emotional negotiation skills into language model agents. EmoDistill decomposes emotional strategy into emotion selection and emotion expression: an Implicit Q-Learning (IQL) selector learns \emph{which} emotion to express, while a Low-Rank Adaptation (LoRA)-based policy learns \emph{how} to express it through Supervised Fine-Tuning (SFT) and Judge Policy Optimization (JPO). Across four emotion-sensitive, high-stakes negotiation domains, SLM policies trained under the EmoDistill framework achieve the highest utility, outperforming vanilla SLM/LLM baselines and IQL-only emotion selection. Ablations show that emotion conditioning is essential, and transfer studies demonstrate generalization across domains, unseen counterparties, and trained-vs-trained tournaments. Overall, EmoDistill learns skills from offline agent-to-agent interactions, avoiding costly online negotiation during training.
- Abstract(参考訳): ポストトレーニング後のLLMは、応答を人間の好みに合わせるように最適化され、安全で丁寧で会話に適している。
しかし、敵対的な交渉においては、このアライメントは脆弱性となる可能性がある。
GoEmotionsをベースとした情動的プロンプトを用いて、感情は交渉結果を大きく変えることを示し、感情は表面的なスタイルではなく戦略的行動チャネルであることを示唆する。
そこで我々は,言語モデルエージェントに感情的交渉スキルを蒸留するオフラインフレームワークである「textbf{EmoDistill}」を紹介した。
Implicit Q-Learning (IQL) セレクタは感情を表現するために \emph{which} を学習し、LoRA (Lo-Rank Adaptation) ベースのポリシーは "emph{how}" を学び、それをSupervised Fine-Tuning (SFT) とジャッジポリシー最適化 (JPO) を通して表現する。
EmoDistillフレームワークでトレーニングされたSLMポリシーは、4つの感情に敏感で高い交渉領域にまたがって、最高のユーティリティを実現し、バニラSLM/LLMベースラインを上回り、IQLのみの感情選択を実現している。
アブレーションは感情条件付けが不可欠であることを示し、移行研究はドメイン、目に見えない相手、訓練されたvs訓練トーナメントをまたいだ一般化を示す。
全体として、EmoDistillはオフラインのエージェント対エージェントのインタラクションからスキルを学び、トレーニング中に高価なオンライン交渉を避ける。
関連論文リスト
- EQ-Negotiator: Dynamic Emotional Personas Empower Small Language Models for Edge-Deployable Credit Negotiation [66.09161596959771]
小型言語モデル (SLM) は実用的な代替手段を提供するが、大規模言語モデル (LLM) と比較して大きな性能差がある。
本稿では,感情的ペルソナを用いて,この能力ギャップを橋渡しする新しいフレームワークであるEQ-Negotiatorを紹介する。
EQ-Negotiator を用いた 7B パラメータ言語モデルは,ベースライン LLM の 10 倍以上の大きさで,債務回復と交渉効率が向上することを示す。
論文 参考訳(メタデータ) (2025-11-05T11:25:07Z) - EvoEmo: Towards Evolved Emotional Policies for Adversarial LLM Agents in Multi-Turn Price Negotiation [61.627248012799704]
既存のLarge Language Models (LLM)エージェントは、そのような交渉における感情の機能的役割をほとんど見落としている。
本稿では,交渉における動的感情表現を最適化する進化的強化学習フレームワークであるEvoEmoを紹介する。
論文 参考訳(メタデータ) (2025-09-04T15:23:58Z) - EmoCAST: Emotional Talking Portrait via Emotive Text Description [56.42674612728354]
EmoCASTは、正確なテキスト駆動感情合成のための拡散ベースのフレームワークである。
外観モデリングでは、感情的なプロンプトはテキスト誘導の分離された感情的モジュールを通して統合される。
EmoCASTは、現実的で感情的に表現され、音声同期されたトーキーヘッドビデオを生成する、最先端のパフォーマンスを実現する。
論文 参考訳(メタデータ) (2025-08-28T10:02:06Z) - EmoDebt: Bayesian-Optimized Emotional Intelligence for Strategic Agent-to-Agent Debt Recovery [65.30120701878582]
大規模言語モデル(LLM)エージェントは、負債収集のような感情に敏感なドメインの悪用に対して脆弱である。
EmoDebtは、ネゴシエーションにおける感情を表現するモデルの能力を、シーケンシャルな意思決定問題として再設計する感情インテリジェンスエンジンである。
EmoDebtは重要な戦略的堅牢性を実現し、非適応性と感情に依存しないベースラインを大幅に上回っている。
論文 参考訳(メタデータ) (2025-03-27T01:41:34Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。