論文の概要: Unbiased Dynamic Multimodal Fusion
- arxiv url: http://arxiv.org/abs/2603.19681v1
- Date: Fri, 20 Mar 2026 06:29:59 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-03-23 19:48:39.025786
- Title: Unbiased Dynamic Multimodal Fusion
- Title(参考訳): ダイナミックマルチモーダル核融合
- Authors: Shicai Wei, Kaijie Zhang, Luyi Chen, Tao He, Guiduo Duan,
- Abstract要約: 従来のマルチモーダル手法は静的なモダリティの品質を前提としており、動的実世界のシナリオにおける適応性を制限している。
モーフィアデータに制御ノイズを付加し,その強度をモーフィア特徴から予測するノイズ認識不確実性推定器を提案する。
これにより、モデルが特徴量と雑音レベルの明確な対応を学習し、低騒音条件と高騒音条件の両方で正確な不確実性の測定を可能にする。
- 参考スコア(独自算出の注目度): 17.111704025557376
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Traditional multimodal methods often assume static modality quality, which limits their adaptability in dynamic real-world scenarios. Thus, dynamical multimodal methods are proposed to assess modality quality and adjust their contribution accordingly. However, they typically rely on empirical metrics, failing to measure the modality quality when noise levels are extremely low or high. Moreover, existing methods usually assume that the initial contribution of each modality is the same, neglecting the intrinsic modality dependency bias. As a result, the modality hard to learn would be doubly penalized, and the performance of dynamical fusion could be inferior to that of static fusion. To address these challenges, we propose the Unbiased Dynamic Multimodal Learning (UDML) framework. Specifically, we introduce a noise-aware uncertainty estimator that adds controlled noise to the modality data and predicts its intensity from the modality feature. This forces the model to learn a clear correspondence between feature corruption and noise level, allowing accurate uncertainty measure across both low- and high-noise conditions. Furthermore, we quantify the inherent modality reliance bias within multimodal networks via modality dropout and incorporate it into the weighting mechanism. This eliminates the dual suppression effect on the hard-to-learn modality. Extensive experiments across diverse multimodal benchmark tasks validate the effectiveness, versatility, and generalizability of the proposed UDML. The code is available at https://github.com/shicaiwei123/UDML.
- Abstract(参考訳): 従来のマルチモーダル手法は静的なモダリティの品質を前提としており、動的実世界のシナリオにおける適応性を制限している。
そこで, 動的マルチモーダル法は, モダリティの品質を評価し, コントリビューションを調整するために提案される。
しかし、それらは一般的に経験的な指標に依存しており、ノイズレベルが極端に低い場合や高い場合、モダリティの品質を測ることに失敗した。
さらに、既存の手法では、各モダリティの初期寄与が同じであると仮定し、本質的なモダリティ依存バイアスを無視する。
その結果、学習し難いモダリティは2倍に罰せられ、動的核融合の性能は静的核融合よりも劣る可能性がある。
これらの課題に対処するため,Unbiased Dynamic Multimodal Learning (UDML) フレームワークを提案する。
具体的には、モーダリティデータに制御ノイズを加え、モーダリティ特徴からその強度を予測するノイズ認識不確実性推定器を提案する。
これにより、モデルが特徴量と雑音レベルの明確な対応を学習し、低騒音条件と高騒音条件の両方で正確な不確実性の測定を可能にする。
さらに,マルチモーダルネットワークにおけるモダリティ依存バイアスをモダリティドロップアウトにより定量化し,重み付け機構に組み込む。
これにより、ハード・トゥ・ラーン・モダリティに対する二重抑制効果が排除される。
多様なマルチモーダルベンチマークタスクにわたる広範囲な実験は、提案したUDMLの有効性、汎用性、および一般化性を検証する。
コードはhttps://github.com/shicaiwei123/UDMLで公開されている。
関連論文リスト
- Test-time Adaptive Hierarchical Co-enhanced Denoising Network for Reliable Multimodal Classification [55.56234913868664]
マルチモーダルデータを用いた信頼性学習のためのTAHCD(Test-time Adaptive Hierarchical Co-enhanced Denoising Network)を提案する。
提案手法は,最先端の信頼性の高いマルチモーダル学習手法と比較して,優れた分類性能,堅牢性,一般化を実現する。
論文 参考訳(メタデータ) (2026-01-12T03:14:12Z) - Multimodal Negative Learning [55.67017420486548]
我々は新しい学習パラダイム"学習すべきでない"(Negative Learning)を提案する。
弱いモダリティのターゲットクラス予測を強化する代わりに、支配的なモダリティは弱いモダリティを動的に導き、非ターゲットクラスを抑える。
これは決定空間を安定化させ、モダリティ固有の情報を保存する。
論文 参考訳(メタデータ) (2025-10-23T11:47:11Z) - BTW: A Non-Parametric Variance Stabilization Framework for Multimodal Model Integration [20.600001069987318]
トレーニング中のモダリティ重要度を動的に調整するために,BTW(Beyond Two-modality Weighting)を提案する。
BTWは、各ユニモーダルと現在のマルチモーダル予測とのばらつきを測定することで、サンプル毎のKL重みを計算する。
本手法は回帰性能と多クラス分類精度を大幅に向上させる。
論文 参考訳(メタデータ) (2025-08-25T23:00:38Z) - Dynamic Modality Scheduling for Multimodal Large Models via Confidence, Uncertainty, and Semantic Consistency [0.0]
本研究では,各モダリティの寄与をサンプル単位で適応的に調整する新しいフレームワークである動的モダリティスケジューリング(DMS)を提案する。
VQA、画像テキスト検索、キャプションタスクの実験結果から、DMSはクリーンとロバストの両方のパフォーマンスを著しく改善することが示された。
論文 参考訳(メタデータ) (2025-06-15T05:15:52Z) - Improving Multimodal Learning Balance and Sufficiency through Data Remixing [14.282792733217653]
弱いモダリティを強制する方法は、単調な充足性とマルチモーダルなバランスを達成できない。
マルチモーダルデータのデカップリングや,各モーダルに対するハードサンプルのフィルタリングなど,モダリティの不均衡を軽減するマルチモーダルデータリミックスを提案する。
提案手法は既存の手法とシームレスに統合され,CREMADでは約6.50%$uparrow$,Kineetic-Soundsでは3.41%$uparrow$の精度が向上する。
論文 参考訳(メタデータ) (2025-06-13T08:01:29Z) - A Study of Dropout-Induced Modality Bias on Robustness to Missing Video
Frames for Audio-Visual Speech Recognition [53.800937914403654]
AVSR(Advanced Audio-Visual Speech Recognition)システムは、欠落したビデオフレームに敏感であることが観察されている。
ビデオモダリティにドロップアウト技術を適用することで、フレーム不足に対するロバスト性が向上する一方、完全なデータ入力を扱う場合、同時に性能損失が発生する。
本稿では,MDA-KD(Multimodal Distribution Approximation with Knowledge Distillation)フレームワークを提案する。
論文 参考訳(メタデータ) (2024-03-07T06:06:55Z) - Cross-Attention is Not Enough: Incongruity-Aware Dynamic Hierarchical
Fusion for Multimodal Affect Recognition [69.32305810128994]
モダリティ間の同調性は、特に認知に影響を及ぼすマルチモーダル融合の課題となる。
本稿では,動的モダリティゲーティング(HCT-DMG)を用いた階層型クロスモーダルトランスを提案する。
HCT-DMG: 1) 従来のマルチモーダルモデルを約0.8Mパラメータで上回り、2) 不整合が認識に影響を及ぼすハードサンプルを認識し、3) 潜在レベルの非整合性をクロスモーダルアテンションで緩和する。
論文 参考訳(メタデータ) (2023-05-23T01:24:15Z) - Trustworthy Multimodal Regression with Mixture of Normal-inverse Gamma
Distributions [91.63716984911278]
このアルゴリズムは、異なるモードの適応的統合の原理における不確かさを効率的に推定し、信頼できる回帰結果を生成する。
実世界のデータと実世界のデータの両方に対する実験結果から,多モード回帰タスクにおける本手法の有効性と信頼性が示された。
論文 参考訳(メタデータ) (2021-11-11T14:28:12Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。