論文の概要: AnthroDial: Benchmarking LLM Anthropomorphism in Autonomous Social Interaction
- arxiv url: http://arxiv.org/abs/2609.37853v1
- Date: Tue, 29 Sep 2026 15:43:07 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-30 21:28:47.701563
- Title: AnthroDial: Benchmarking LLM Anthropomorphism in Autonomous Social Interaction
- Title(参考訳): 自律的社会的相互作用におけるLLM擬人化のベンチマーク
- Abstract要約: AnthroDialは、オープンエンドインタラクションにおいて、信頼できる人間のようなソーシャルエージェントを開発するためのフレームワークである。
MindFlowは動的Mind Bufferを通じて、自律的で非同期で適応的なコミュニケーションを可能にする。
CAPS-Evalは、人為的相互作用の認知的、感情的、行動的次元を評価するための理論的な基盤となるフレームワークである。
- 参考スコア(独自算出の注目度): 16.22592455446366
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Large language models (LLMs) are increasingly deployed as social agents, yet credible human-like interaction requires more than fluent responses or persona consistency. Agents must autonomously decide whether, when, and how to communicate while adapting to evolving contexts, goals, and relationships. Existing research, however, lacks a unified approach to enabling, evaluating, and improving such capabilities in continuous, open-ended interaction. We introduce AnthroDial, a unified framework for developing anthropomorphic social agents from three complementary aspects: MindFlow, a lightweight interaction harness that enables autonomous, asynchronous, and adaptive communication through a dynamic Mind Buffer; CAPS-Eval, a theory-grounded framework for evaluating cognitive, affective, and behavioral dimensions of anthropomorphic interaction; and a scalable training paradigm that combines SEEDS for environment expansion with DiAPO for adaptive capability optimization. We further construct evaluation datasets covering everyday communication, game interaction, and long-horizon character interaction. Extensive experiments across diverse models and scenarios demonstrate improved interaction autonomy and naturalness, validate the reliability, discriminativeness, and agreement with human rankings of CAPS-Eval, and confirm the effectiveness of our training paradigm. Together, these components provide a unified framework for developing credible human-like social agents in open-ended interaction.
- Abstract(参考訳): 大規模言語モデル(LLM)は、社会的エージェントとしてますますデプロイされるが、信頼できる人間のような相互作用には、流動的な応答やペルソナの一貫性以上のものが必要である。
エージェントは、進化するコンテキスト、目標、関係に適応しながら、いつ、どのようにコミュニケーションするかを自律的に決定する必要がある。
しかし、既存の研究は、継続的かつオープンな相互作用においてそのような機能を実現し、評価し、改善するための統一的なアプローチを欠いている。
我々は,3つの相補的な側面から人為的社会エージェントを開発するための統一的なフレームワーク,MindFlow,動的マインドバッファによる自律的・非同期的・適応的なコミュニケーションを可能にする軽量なインタラクションハーネス,CAPS-Eval,人為的相互作用の認知的・情緒的・行動的次元を評価するための理論的基礎的フレームワーク,環境拡張のためのSEEDSとDiAPOを組み合わせたスケーラブルなトレーニングパラダイムを紹介する。
さらに,日常コミュニケーション,ゲームインタラクション,長軸キャラクタインタラクションを含む評価データセットを構築した。
多様なモデルやシナリオにまたがる広範な実験により、相互作用の自律性と自然性が向上し、信頼性、差別性、およびCAPS-Evalの人間のランクとの一致が検証され、トレーニングパラダイムの有効性が確認された。
これらのコンポーネントは、オープンエンドインタラクションにおいて、信頼できる人間のようなソーシャルエージェントを開発するための統一的なフレームワークを提供する。
関連論文リスト
- PersLLM: A Personified Training Approach for Large Language Models [66.16513246245401]
データ構築とモデルチューニングを改善するためのフレームワークPersLLMを提案する。
データ利用が不十分な場合には、Chain-of-Thoughtプロンプトやアンチインダクションといった戦略を取り入れます。
厳密な振舞いパターンを設計し,モデルの性格の特異性とダイナミズムを高めるために自動DPOを導入する。
論文 参考訳(メタデータ) (2024-07-17T08:13:22Z) - AntEval: Evaluation of Social Interaction Competencies in LLM-Driven
Agents [65.16893197330589]
大規模言語モデル(LLM)は、幅広いシナリオで人間の振る舞いを再現する能力を示した。
しかし、複雑なマルチ文字のソーシャルインタラクションを扱う能力については、まだ完全には研究されていない。
本稿では,新しいインタラクションフレームワークと評価手法を含むマルチエージェントインタラクション評価フレームワーク(AntEval)を紹介する。
論文 参考訳(メタデータ) (2024-01-12T11:18:00Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。