論文の概要: VoiceNet: Fine-Grained Voice Understanding Beyond Emotion at Scale
- arxiv url: http://arxiv.org/abs/2609.32016v1
- Date: Fri, 25 Sep 2026 21:28:16 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-08 03:50:08.284349
- Title: VoiceNet: Fine-Grained Voice Understanding Beyond Emotion at Scale
- Title(参考訳): VoiceNet: スケールの感情を超えた細分化された音声理解
- Abstract要約: 本稿では,人間の声質評価基準であるVoiceNetを紹介する。
Emilia corpus の感情アノテート版である Emolia もリリースされ,高密度な MOSS-Audio Thinking アノテーションによって強化された再バランスサブセットが提供される。
- 参考スコア(独自算出の注目度): 29.84704827157091
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Expressive speech synthesis has outpaced expressive speech perception: systems now render fine-grained vocal performances that no public benchmark can score. Most benchmarks for this inverse problem stop at six to nine basic emotion categories, largely on acted speech. This paper introduces VoiceNet, a human-annotated representation-level benchmark for voice performance understanding on permissively-licensed in-the-wild speech. VoiceNet has two subsets: VoiceNet-Emo applies a 40-emotion taxonomy with three expert ratings per item, and VoiceNet-Ext, a preliminary subset, scores 57 talking-style attributes including speaking rate, vocal tension, breathiness, and register. The paper also releases Emolia, an emotion-annotated version of the Emilia corpus, with a curated rebalanced subset enriched by dense MOSS-Audio Thinking annotations. Two voice-text contrastive baselines train on this data: a 110M-parameter VoiceCLAP-Small for fast large-scale data filtering and a 7B VoiceCLAP-Large for state-of-the-art performance. Both outperform existing CLAP baselines, which sit near chance on VoiceNet-Emo. On VoiceNet-Emo, VoiceCLAP-Large aligns more closely with the aggregate expert consensus than individual experts agree with one another: a comparison against the majority label rather than evidence of surpassing human emotion perception. All systems evaluated here are voice-text embedding models: VoiceNet scores representation-level attribute recognition and retrieval, not end-to-end spoken-dialogue behaviour. Clustering and filtering uncurated speech corpora into subsets that span diverse talking styles and emotions remains an open challenge; VoiceCLAP embeddings offer a promising tool for this task. VoiceNet, Emolia, and VoiceCLAP are publicly available for research use.
- Abstract(参考訳): 表現型音声合成は、表現型音声知覚よりも優れており、システムは、公的なベンチマークが採点できないような、きめ細かい発声を処理している。
この逆問題に対するほとんどのベンチマークは、6から9つの基本的な感情カテゴリーで停止する。
本稿では,人間の声質評価基準であるVoiceNetを紹介する。
VoiceNetには2つのサブセットがあり、VoiceNet-Emoはアイテムごとに3つの専門家評価を持つ40の感情分類を適用し、VoiceNet-Extは予備サブセットであり、発話率、声の緊張、呼吸性、レジスターを含む57のトーキングスタイルの属性をスコア付けする。
この論文は、Emilia corpusの感情アノテートバージョンであるEmoliaもリリースしている。
高速な大規模データフィルタリングのための110MパラメータVoiceCLAP-Smallと、最先端のパフォーマンスのための7BVoiceCLAP-Largeである。
どちらも既存のCLAPベースラインを上回り、VoiceNet-Emoでほぼチャンスに迫られている。
VoiceNet-Emoでは、VoiceCLAP-Largeは、個人の専門家同士が同意するよりも、専門家の集合的な合意とより密接に一致している。
ボイスネットは、エンドツーエンドの音声対話行動ではなく、表現レベルの属性認識と検索をスコアする。
さまざまな話し方や感情にまたがるサブセットに、未処理の音声コーパスをクラスタ化してフィルタリングすることは、依然としてオープンな課題である。
VoiceNet、Emolia、VoiceCLAPは研究用に公開されている。
関連論文リスト
- RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems [5.434878413067881]
テキスト音声(TTS)、音声音声(STS)、音声理解(SU)、音声認識(ASR)による音声AI評価のための多次元ベンチマークであるReal World Voice EQ Benchを紹介する。
評価の結果, 性能は高次元比であることが示唆された。
論文 参考訳(メタデータ) (2026-07-16T11:15:16Z) - UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models [37.12134339309316]
我々は,複数のきめ細かい音声スタイル制御のために開発された,最初の大規模音声対話データセットであるUltraVoiceを紹介する。
SLAM-OmniやVocalNet on UltraVoiceのような微調整型の先行モデルは、その微調整性を大幅に向上させる。
URO-Benchベンチマークでは、微調整されたモデルでは、コア理解、推論、会話能力が大幅に向上した。
論文 参考訳(メタデータ) (2025-10-26T09:06:55Z) - MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions [70.93364531054273]
音声と視覚を融合させる音声アシスタントの能力を評価する最初のベンチマークであるMultiVoxを紹介する。
具体的には、MultiVoxには、多種多様なパラ言語的特徴を包含する1000の人間の注釈付き音声対話が含まれている。
10の最先端モデルに対する我々の評価は、人間はこれらのタスクに長けているが、現在のモデルは、常に文脈的に接地された応答を生成するのに苦労していることを示している。
論文 参考訳(メタデータ) (2025-07-14T23:20:42Z) - EmoNet-Voice: A Fine-Grained, Expert-Verified Benchmark for Speech Emotion Detection [19.43600992826571]
本稿では,音声感情検出のための新しいリソースであるEmoNet-Voiceを紹介する。
EmoNet-Voiceは、40の感情カテゴリーの細かいスペクトルでSERモデルを評価するように設計されている。
また、人間の専門家と高い合意を得て、音声感情認識の新しい標準となる共感型Insight Voiceモデルも導入する。
論文 参考訳(メタデータ) (2025-06-11T15:06:59Z) - EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting [48.56693150755667]
感情制御可能な新しいTSモデルであるEmoVoiceを提案する。
EmoVoiceは、大きな言語モデル(LLM)を利用して、きめ細かいフリースタイルの自然言語感情制御を可能にする。
EmoVoiceは、英語のEmoVoice-DBテストセットで最先端のパフォーマンスを達成する。
論文 参考訳(メタデータ) (2025-04-17T11:50:04Z) - Time out of Mind: Generating Rate of Speech conditioned on emotion and
speaker [0.0]
感情によって条件付けされたGANをトレーニングし、与えられた入力テキストに価値ある長さを生成する。
これらの単語長は相対的中性音声であり、テキスト音声システムに提供され、より表現力のある音声を生成する。
我々は,中性音声に対する客観的尺度の精度向上と,アウト・オブ・ボックスモデルと比較した場合の幸福音声に対する時間アライメントの改善を実現した。
論文 参考訳(メタデータ) (2023-01-29T02:58:01Z) - Limited Data Emotional Voice Conversion Leveraging Text-to-Speech:
Two-stage Sequence-to-Sequence Training [91.95855310211176]
感情的音声変換は、言語内容と話者のアイデンティティを保ちながら、発話の感情状態を変えることを目的としている。
本研究では,感情音声データ量の少ない連続音声変換のための新しい2段階学習戦略を提案する。
提案フレームワークはスペクトル変換と韻律変換の両方が可能であり、客観的評価と主観評価の両方において最先端のベースラインを大幅に改善する。
論文 参考訳(メタデータ) (2021-03-31T04:56:14Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。