論文の概要: Disclosure-Gated User Simulation for Companion-Agent Evaluation
- arxiv url: http://arxiv.org/abs/2609.00982v1
- Date: Tue, 01 Sep 2026 09:39:41 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-02 16:31:36.522317
- Title: Disclosure-Gated User Simulation for Companion-Agent Evaluation
- Title(参考訳): コンパニオン・エージェント評価のための開示型ユーザシミュレーション
- Authors: Yao Liu, Yu He,
- Abstract要約: 私たちは大きな言語モデルを使ってユーザをプレイし、その仕様に対してユーザシミュレータをトレーニングします。
ゲーティング行動は、トレーニングコーパスの合成ブランチから学習され、実際のブランチは、人々が話す方法と反応する方法を提供する。
ランク付けは順序保存でなければならず、絶対スコアはスケール安定でなければなりません。
- 参考スコア(独自算出の注目度): 14.962552394114894
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Using a large language model to play the user is now standard in scalable evaluation. It has a repeatedly diagnosed failure: the simulated user is excessively cooperative, so a system under test can score by the sheer number of questions it asks rather than by making the user willing to speak. We answer with a disclosure gate conditioning information release on the companion agent's behaviour: its state is a ladder of five ordered gates, merged onto three observable depth layers. We specify, ablate, and audit it, and train a user simulator against that specification. Gating behaviour is learned from the training corpus's synthetic branch, while the real branch supplies how people speak and react; after training, the simulator need not be told at runtime which gate each item sits behind. The gate is a load-bearing component of the environment: on the English corpus of a published companion-agent benchmark (CompanionBench), once training no longer states per example which gate each item sits behind, the largest rank displacement across 12 systems under test exceeds the noise band set by re-running that environment under a new seed, while per-system scores show no detectable change. We state two acceptance criteria: a ranking must be order-preserving, and absolute scores must be scale-stable. Of the candidates we examine, only one passes both -- the simulator we release -- and its leaderboard correlates at 0.993 with the benchmark's original simulator. By contrast, prompting a frontier model as the simulator barely moves the ranking while shifting every score upward -- a shift invisible to anyone checking the ranking alone. The environment we specify is the one that benchmark already used. That publication describes the mechanism in about four hundred words, and we supply what it lacked: specification, ablations, human studies, negative controls, and downstream sensitivity analysis.
- Abstract(参考訳): 大規模言語モデルを使用してユーザをプレイすることが、スケーラブルな評価において標準になった。
シミュレーションされたユーザは過度に協力的であり、テスト対象のシステムは、ユーザが話す意思を持たせるのではなく、質問の数によってスコアを付けることができる。
その状態は5つの順序付けられたゲートのはしごであり、3つの観測可能な深さ層にマージされる。
私たちはそれを指定し、修正し、監査し、その仕様に対してユーザーシミュレータを訓練します。
ゲーティングの動作は、トレーニングコーパスの合成ブランチから学習され、実際のブランチは、人々が話し、反応する方法を提供する。
ゲートは環境の負荷を伴う要素である: 発行されたコンパニオンエージェントベンチマーク(CompanionBench)の英語コーパスでは、各項目が背後にある場合の1つ当たりのトレーニングが終了すると、テスト対象の12システムにおける最大のランク変位は、その環境を新しいシードの下で再実行することで設定されたノイズバンドを超え、システムスコアは検出可能な変化を示さない。
ランク付けは順序保存でなければならず、絶対スコアはスケール安定でなければなりません。
調査対象のうち、リリースしたシミュレーターの2つをパスしたのは1つだけであり、そのリーダーボードはベンチマークのオリジナルのシミュレーターと0.993で相関している。
対照的に、シミュレーターがすべてのスコアを上向きにシフトしながらランキングをほとんど動かさないように、フロンティアモデルを推進します。
私たちが指定する環境は、すでにベンチマークが使用しています。
その出版物では、このメカニズムを約400語で記述し、仕様、改善、人間研究、ネガティブコントロール、下流の感度分析など、欠けているものを提供しています。
関連論文リスト
- CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship [7.643253506288442]
既存のベンチマークでは、手書きのシナリオを使用し、シミュレーターを誘導する。
対話型バイリンガルベンチマークCompanionBenchを紹介する。
実際のデータにシナリオとトレーニングされたユーザーシミュレータの両方を基盤とした最初のコンパニオンベンチマークである。
論文 参考訳(メタデータ) (2026-08-03T10:44:04Z) - Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents [28.85410475505536]
Persona Policies (PPol) は、ユーザシミュレータの現実的な振る舞い変化を誘導するコントロール層である。
PPolはドメイン内のあらゆるタスクに対して、多種多様な人間のようなペルソナを生産する。
PPolで訓練されたエージェントは、挑戦的で非分配的な振る舞いに対してより堅牢であり、既存のシミュレートされたインタラクションのみに関するトレーニングに比べて、タスクの成功率が+17%向上する。
論文 参考訳(メタデータ) (2026-05-13T02:16:51Z) - SimEval-IR: A Unified Toolkit and Benchmark Suite for Evaluating User Simulators and Search Sessions [1.1105673928718571]
オープンソースのツールキットとベンチマークスイートであるSimEval-IRについて述べる。
SimEval-IR は,(1) セッション検索と対話を統一する標準セッションスキーマ,(2) 行動リアリズム,RATE スタイルの推定によるテスタの信頼性,および2つの言語と4つのシミュレーターファミリーの4つの実データセットのベースライン結果に関する3つのベンチマークを提供する。
論文 参考訳(メタデータ) (2026-04-30T13:56:18Z) - SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants? [61.07963107032645]
大規模言語モデル(LLM)は、対話型アプリケーションでますます使われている。
人間の評価は、マルチターン会話におけるパフォーマンスを評価するためのゴールドスタンダードのままである。
我々は、909の注釈付き人間とLLMの会話を2つの対話タスクで行うベンチマークであるSimulatorArenaを紹介した。
論文 参考訳(メタデータ) (2025-10-06T23:17:44Z) - Metaphorical User Simulators for Evaluating Task-oriented Dialogue
Systems [80.77917437785773]
タスク指向対話システム(TDS)は、主にオフラインまたは人間による評価によって評価される。
本稿では,エンド・ツー・エンドのTDS評価のためのメタファ型ユーザシミュレータを提案する。
また,異なる機能を持つ対話システムなどの変種を生成するためのテスタベースの評価フレームワークを提案する。
論文 参考訳(メタデータ) (2022-04-02T05:11:03Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。