論文の概要: MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion
- arxiv url: http://arxiv.org/abs/2609.40087v1
- Date: Wed, 30 Sep 2026 16:31:01 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-01 18:57:28.068871
- Title: MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion
- Title(参考訳): MeanVoiceFlow2: 高速ワンステップゼロショット音声変換のための平均フローとコンテンツエンコーダの併用最適化
- Abstract要約: MeanVoiceFlow2はフローベースの変換モジュールと計算効率の良いコンテントエンコーダを共同で最適化するフレームワークである。
このモデルは、MeanVoiceFlowを用いた変換蒸留と実データの再構成によって訓練される。
ゼロショットVCの実験では、MeanVoiceFlow2は知覚品質が高く、MeanVoiceFlowよりも約9倍高速な推論を実現した。
- 参考スコア(独自算出の注目度): 42.55959060773461
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Flow-matching approaches to voice conversion (VC) have gained attention owing to their high speech quality and strong speaker similarity. Among them, one-step models such as MeanVoiceFlow are particularly attractive because they enable efficient inference; however, their reliance on a computationally intensive content encoder remains a bottleneck. We therefore propose MeanVoiceFlow2, a framework that jointly optimizes a flow-based conversion module and a computationally efficient content encoder. The model is trained through conversion distillation using MeanVoiceFlow and the reconstruction of real data. We further incorporate diffusion-GAN training with sample mixing and teacher-guided conditioning augmentation to enhance realism and disentanglement. Experiments on zero-shot VC showed that MeanVoiceFlow2 achieved higher perceptual quality and approximately $9\times$ faster inference than MeanVoiceFlow while maintaining comparable speaker similarity. Audio samples are available at https://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/meanvoiceflow2/.
- Abstract(参考訳): 音声変換(VC)に対するフローマッチング手法は,高い音声品質と強い話者類似性から注目されている。
中でもMeanVoiceFlowのようなワンステップモデルは、効率的な推論を可能にするため特に魅力的だが、計算集約的なコンテントエンコーダへの依存は依然としてボトルネックとなっている。
そこで我々は,フローベース変換モジュールと計算効率の良いコンテントエンコーダを共同で最適化するフレームワークであるMeanVoiceFlow2を提案する。
このモデルは、MeanVoiceFlowを用いた変換蒸留と実データの再構成によって訓練される。
さらに,拡散GANトレーニングにサンプルミキシングと教師誘導条件強化を併用し,現実性や絡み合いを高めた。
ゼロショットVCの実験では、MeanVoiceFlow2は知覚品質が高く、MeanVoiceFlowよりも約9\times$高速な推論を実現した。
オーディオサンプルはhttps://www.kecl.ntt.co.jp/people/ Kaneko.takuhiro/projects/meanvoiceflow2/で入手できる。
関連論文リスト
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space [68.11390597559101]
X-VCはゼロショットストリーミングVCシステムであり、事前訓練されたニューラルネットワークの潜在空間でワンステップ変換を行う。
X-VCは、英語と中国語の両方で最高のストリーミングWERを達成する。
論文 参考訳(メタデータ) (2026-04-14T08:42:10Z) - Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching [51.70360630470263]
Video-to-audio (V2A) は、サイレントビデオからコンテンツマッチング音声を合成することを目的としている。
本稿では,修正フローマッチングに基づくV2AモデルであるFrierenを提案する。
実験により、フリーレンは世代品質と時間的アライメントの両方で最先端のパフォーマンスを達成することが示された。
論文 参考訳(メタデータ) (2024-06-01T06:40:22Z) - VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching [14.7974342537458]
VoiceFlowは,修正フローマッチングアルゴリズムを用いて,限られたサンプリングステップ数で高い合成品質を実現する音響モデルである。
単話者コーパスと多話者コーパスの主観的および客観的評価の結果,VoiceFlowの合成品質は拡散コーパスに比べて優れていた。
論文 参考訳(メタデータ) (2023-09-10T13:47:39Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。