論文の概要: ResDiffFRG: Residual Diffusion for Multiple Appropriate Facial Reaction Generation
- arxiv url: http://arxiv.org/abs/2609.33749v1
- Date: Sun, 27 Sep 2026 16:50:05 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-04 17:04:08.929147
- Title: ResDiffFRG: Residual Diffusion for Multiple Appropriate Facial Reaction Generation
- Title(参考訳): ResDiffFRG:複数顔反応生成のための残留拡散
- Abstract要約: 人間の話者とリスナーの会話では、リスナーの表情反応により、リスナーの感情状態が正確に知覚される。
人間の顔反応は非決定論的であるため、現実的な人間とエージェントの相互作用には、複数の適切な人間の顔反応を生成する能力が不可欠である。
ResDiffFRGは拡散過程を話者行動に固定する新しい拡散型MAFRGフレームワークである。
- 参考スコア(独自算出の注目度): 26.89886895158014
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: In dyadic human speaker-listener conversations, the listener's facial reactions allows the speaker to accurately perceive the listener's emotional states. Since human facial reactions are non-deterministic, the ability to generate multiple appropriate human-like facial reactions is crucial for realistic human-agent interactions. Although diffusion models are naturally suited to such one-to-many generation, existing diffusion-based Multiple Appropriate Facial Reaction Generation (MAFRG) methods attempt to denoise random Gaussian initialisations directly into multiple appropriate facial reactions (AFRs). These random initialisations are usually not well-aligned with the target listener facial reaction, which requires complex denoising trajectories from these initialisations, and subsequently creates substantial opportunities for deviations away from the range of trajectories leading to appropriate AFRs. Given the inherent mimicry between the human listener's and speaker's facial behaviours, we address the above denoising trajectory issue by leveraging this strong prior. Specifically, we propose ResDiffFRG, a novel diffusion-based MAFRG framework that explicitly anchors the diffusion process to the speaker behaviour by defining its diffusion target as the residual between the speaker anchor and an AFR. The denoiser only needs to model the comparatively small, reaction-specific residual needed to transform this anchor into an AFR, rather than reconstructing the complete reaction from an unstructured state. Extensive experiments show that ResDiffFRG achieves large improvements in correlation-based appropriateness over existing methods. Our denoising trajectory analysis showed that even at the start of the denoising trajectory, ResDiffFRG already achieves a higher facial-reaction correlation score than the Gaussian Diffusion baseline does after completing 60% of its denoising trajectory.
- Abstract(参考訳): 人間の話者とリスナーの会話では、リスナーの表情反応により、リスナーの感情状態が正確に知覚される。
人間の顔反応は非決定論的であるため、現実的な人間とエージェントの相互作用には、複数の適切な人間の顔反応を生成する能力が不可欠である。
拡散モデルは1対多生成に自然に適しているが、既存の拡散に基づく多重顔反応生成(MAFRG)法は、ランダムなガウス初期化を複数の適切な顔反応(AFR)に直接分解しようとする。
これらのランダムな初期化は、通常、ターゲットリスナーの顔反応とうまく一致しないため、これらの初期化から複雑な認知軌道が必要となり、その後、適切なAFRにつながる軌道の範囲から逸脱するかなりの機会が生じる。
人間の聞き手と話者の顔行動に固有の模倣を考慮し,この強みを生かして,上記の認知的軌道問題に対処する。
具体的には、話者アンカーとAFRの間の残差として拡散目標を定義することにより、拡散過程を話者行動に明示的にアンカーする新しい拡散ベースMAFRGフレームワークであるResDiffFRGを提案する。
デノイザーは、このアンカーを非構造状態から完全な反応を再構築するのではなく、AFRに変換するのに必要な比較的小さな反応特異的残基をモデル化するのみである。
大規模な実験により、ResDiffFRGは既存の手法よりも相関に基づく適切性を大幅に向上することが示された。
ResDiffFRGは難聴軌跡の60%を完了した後のガウス拡散ベースラインよりも高い顔面反応相関スコアをすでに達成している。
関連論文リスト
- PhoenixSR: Generative Heterogeneous Distillation Unleashes Efficient Models for Real-World Super-Resolution [77.62046566633215]
実世界の超解像(SR)は、複雑な低分解能観測から知覚的に現実的な高分解能像を復元する必要がある。
PhoenixSRは独立に設計されたフィードフォワードSRネットワークに拡散先行を転送する生成ヘテロジニアス蒸留フレームワークである。
3つのSRベンチマークと6つのフィードフォワードバックボーン(SwinIR、HAT、Real-ESRGAN、SeeeMoRe)の実験は、ほとんど保存された復元フィデリティと一貫した知覚的改善を示している。
論文 参考訳(メタデータ) (2026-09-25T08:34:32Z) - Uniform Diffusion Models Revisited: Leave-One-Out Denoiser and Absorbing State Reformulation [72.05592785529404]
UDMの標準プラグインブリッジパラメタライゼーションは,後側頭蓋に最適化されないことを示す。
また,UDM関節法をマスク拡散様サンプリング操作に分解して保存する均一拡散の吸収状態再構成も導入した。
論文 参考訳(メタデータ) (2026-05-21T17:27:19Z) - Taming a Retrieval Framework to Read Images in Humanlike Manner for Augmenting Generation of MLLMs [23.638717678491986]
マルチモーダルな大言語モデル(MLLM)は、細粒度の視覚的質問応答でしばしば失敗する。
HuLiRAG(Human-like Retrieval-Augmented Generation)は、マルチモーダル推論を「何」のカスケードとしてステージングするフレームワークである。
論文 参考訳(メタデータ) (2025-10-12T03:22:33Z) - ReactDiff: Fundamental Multiple Appropriate Facial Reaction Diffusion Model [0.9786690381850356]
多様な顔反応を生成するための新しい時間拡散フレームワークであるReactDiffを提案する。
私たちの重要な洞察は、もっともらしい人間の反応は、時間の経過とともに滑らかさとコヒーレンスを示すということです。
提案手法は, 最先端の反応品質を達成し, 多様性と反応適性に優れる。
論文 参考訳(メタデータ) (2025-10-06T11:30:40Z) - ReactDiff: Latent Diffusion for Facial Reaction Generation [15.490774894749277]
話者の音声・視覚的クリップを考えると、顔反応生成はリスナーの顔反応を予測することを目的としている。
本稿では,多モード変換器と条件拡散を統合した顔反応拡散(ReactDiff)フレームワークを提案する。
実験の結果、ReactDiffは既存のアプローチよりも大幅に優れており、顔反応の相関は0.26、多様性のスコアは0.094である。
論文 参考訳(メタデータ) (2025-05-20T10:01:37Z) - SMRD: SURE-based Robust MRI Reconstruction with Diffusion Models [76.43625653814911]
拡散モデルは、高い試料品質のため、MRIの再生を加速するために人気を博している。
推論時に柔軟にフォワードモデルを組み込んだまま、効果的にリッチなデータプリエントとして機能することができる。
拡散モデル(SMRD)を用いたSUREに基づくMRI再構成を導入し,テスト時の堅牢性を向上する。
論文 参考訳(メタデータ) (2023-10-03T05:05:35Z) - Reconstructing Graph Diffusion History from a Single Snapshot [87.20550495678907]
A single SnapsHot (DASH) から拡散履歴を再構築するための新しいバリセンターの定式化を提案する。
本研究では,拡散パラメータ推定のNP硬度により,拡散パラメータの推定誤差が避けられないことを証明する。
また、DITTO(Diffusion hitting Times with Optimal proposal)という効果的な解法も開発している。
論文 参考訳(メタデータ) (2023-06-01T09:39:32Z) - DiffusionAD: Norm-guided One-step Denoising Diffusion for Anomaly Detection [80.20339155618612]
DiffusionADは、再構成サブネットワークとセグメンテーションサブネットワークからなる、新しい異常検出パイプラインである。
高速なワンステップデノゲーションパラダイムは、同等の再現品質を維持しながら、数百倍の加速を達成する。
異常の出現の多様性を考慮し、複数のノイズスケールの利点を統合するためのノルム誘導パラダイムを提案する。
論文 参考訳(メタデータ) (2023-03-15T16:14:06Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。