論文の概要: MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents
- arxiv url: http://arxiv.org/abs/2609.20152v1
- Date: Thu, 17 Sep 2026 12:44:03 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-20 08:55:54.274743
- Title: MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents
- Title(参考訳): MTVA-Bench:カスケード音声エージェント内の言語モデルの評価
- Abstract要約: そこで, MTVA-Bench (Multi-Turn Voice Agent Benchmark) を提案する。
ベンチマークには49のエージェントが含まれ、レビューされた490のシナリオで動作し、7つの言語をサポートする。
7つのモデルによる調査では、6つのモデルがそれぞれの6.4ポイント以内で正しいツールを選択するが、全体的なスコアは24.4ポイントである。
- 参考スコア(独自算出の注目度): 0.7611870296994722
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Generally, most voice agents are cascaded systems, i.e., an ASR model transcribes the caller's audio, a language model reads the transcript and decides what to say and which backend tools to call, and a TTS model speaks the reply. Nearly all of the decision making happens in the language model, but existing evaluations measure it either too broadly or too narrowly. End-to-end voice benchmarks score the full pipeline, so recognition errors and model errors mix into a single number. LLM benchmarks isolate the model but they do not evaluate what makes real phone calls hard, such as transcription issues, caller's voice being split across messages and the requirement that replies follow the language and script specified. We introduce the Multi-Turn Voice Agent Benchmark (MTVA-Bench), which evaluates the language model on the same conditions it faces inside a cascaded system. The caller is played by an LLM following a set of rubrics and tool calls are answered by a mock backend which responds to the arguments the model actually sent. The benchmark contains 49 agents working across 490 reviewed scenarios and supports 7 languages. Scoring is a combination of deterministic checks on tool calls with two LLM judges, one that scores scenario specific rules and one that grades conversation quality without access to the task. Both judges must cite specific messages from the transcript. Task and conversation scores are weighted equally, since a call can complete its task and still go badly for the caller. In a seven-model study, six of the models select the correct tool within 6.4 points of one another, but their overall scores span 24.4 points. Most of the gap comes from argument values, action ordering, rule compliance, and what the model says around its tool calls.
- Abstract(参考訳): 一般的に、ほとんどの音声エージェントはカスケードシステムであり、すなわち、ASRモデルは発信者の音声を書き起こし、言語モデルは書き起こしを読み、何とどのバックエンドツールを呼び出すかを判断し、TSモデルは応答を話す。
決定のほとんどすべてが言語モデルで行われるが、既存の評価では広すぎるか狭すぎるかが測定されている。
エンドツーエンドの音声ベンチマークが全パイプラインをスコアするので、認識エラーとモデルエラーが1つの番号に混在する。
LLMベンチマークは、モデルを分離するが、文字起こしの問題、発信者の音声がメッセージ間で分割されること、応答が言語やスクリプトに従わなければならないことなど、実際の電話の呼び出しを難しくするものではない。
そこで, MTVA-Bench (Multi-Turn Voice Agent Benchmark) を提案する。
呼び出しは、一連のルーリックに従ってLLMによって実行され、ツールコールはモックバックエンドによって応答され、モデルが実際に送信した引数に応答する。
ベンチマークには49のエージェントが含まれ、レビューされた490のシナリオで動作し、7つの言語をサポートする。
Scoringは、ツールコールに対する決定論的チェックと、シナリオ固有のルールをスコアする2つのLLMジャッジ、タスクにアクセスせずに会話品質をグレードする1つの組み合わせである。
双方の裁判官は、書き起こしから特定のメッセージを引用しなければならない。
タスクと会話のスコアは均等に重み付けされる。
7つのモデルによる調査では、6つのモデルがそれぞれの6.4ポイント以内で正しいツールを選択するが、全体的なスコアは24.4ポイントである。
ギャップのほとんどは、引数値、アクションの順序付け、ルールのコンプライアンス、そしてそのツール呼び出しに関するモデルが言うことに由来する。
関連論文リスト
- DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units [15.409370828694563]
DiscoPhonは、離散音声単位から教師なし音素発見を評価するためのベンチマークである。
6つの開発言語と6つのテスト言語をカバーし、幅広い音韻のコントラストをカバーしている。
論文 参考訳(メタデータ) (2026-03-19T08:31:58Z) - VoiceAgentBench: Are Voice Assistants ready for agentic tasks? [5.639970295197759]
本稿では,現実的な音声エージェント設定におけるSpeechLMの評価ベンチマークであるVoiceAgentBenchを紹介する。
インドの文脈に根ざした5,500以上の合成音声クエリで構成されている。
ツール選択の正確性、構造的整合性、ツールの実行の正しさを測定する。
論文 参考訳(メタデータ) (2025-10-09T09:11:38Z) - AHELM: A Holistic Evaluation of Audio-Language Models [78.20477815156484]
マルチモーダルオーディオ言語モデル(ALM)は、インターリーブされた音声とテキストを入力および出力テキストとして取り込む。
AHELMは、PARADEとCoRe-Benchと呼ばれる2つの新しい合成オーディオテキストデータセットを含む、さまざまなデータセットを集約するベンチマークである。
また、モデル間の等価比較を確保するために、プロンプト、推論パラメータ、評価指標を標準化する。
論文 参考訳(メタデータ) (2025-08-29T07:40:39Z) - Retrieval Augmented End-to-End Spoken Dialog Models [20.896330994089283]
音声信号から直接ダイアログ状態が推測される音声対話アプリケーションにSLMを適用する。
RAG(retrieval-augmented generation)パラダイムにヒントを得て,この弱点を克服する検索拡張SLM(ReSLM)を提案する。
音声MultipleWozタスク(DSTC-11チャレンジ)を用いてReSLMを評価し,この検索によりモデル性能が向上することを確認した。
論文 参考訳(メタデータ) (2024-02-02T18:23:09Z) - SpokenWOZ: A Large-Scale Speech-Text Benchmark for Spoken Task-Oriented Dialogue Agents [70.08842857515141]
SpokenWOZは音声TODのための大規模音声テキストデータセットである。
SpokenWOZでは、クロスターンスロットと推論スロット検出が新たな課題である。
論文 参考訳(メタデータ) (2023-05-22T13:47:51Z) - Towards Zero-shot Learning for Automatic Phonemic Transcription [82.9910512414173]
より難しい問題は、トレーニングデータをゼロにする言語のための音素変換器を構築することだ。
我々のモデルは、トレーニングデータなしで、ターゲット言語で見知らぬ音素を認識できる。
標準的な多言語モデルよりも平均して7.7%の音素誤り率を実現している。
論文 参考訳(メタデータ) (2020-02-26T20:38:42Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。