論文の概要: Do More Modalities Always Help? A Geometric Perspective on Missing-Modality Robustness
- arxiv url: http://arxiv.org/abs/2610.04792v1
- Date: Sat, 03 Oct 2026 22:34:19 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-11 08:53:03.797899
- Title: Do More Modalities Always Help? A Geometric Perspective on Missing-Modality Robustness
- Title(参考訳): モダリティは常に役に立つか? モダリティの欠如を測る幾何学的視点
- Abstract要約: 完全なモダリティで訓練されたモデルは、推論時に1つのモダリティが欠落した場合に、アンモダリティモデルに過小評価される。
このような分解は、主パラメータ部分空間における学習されたクロスモーダル依存関係と密接に関連している。
本稿では,Grassmann的部分空間幾何を構造的部分空間補正に用いる軽量パラメータ編集法を提案する。
- 参考スコア(独自算出の注目度): 55.50054865082731
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Missing modality remains a longstanding challenge in multimodal learning. Existing methods typically address this issue through modality recovery or adaptive strategies. However, they overlook models' internal cross-modal dependencies formed during multimodal training, which later impair robustness. We systematically characterize a counterintuitive deployment-time failure mode: models trained on full modalities can underperform unimodal models when one modality is missing at inference time. This pattern appears across diverse architectures, such as fusion models, CLIP-style two-tower models, and vision-language models. We show that such degradation is closely associated with learned cross-modal dependencies in the principal parameter subspaces. Multimodal training induces structured rotations of these subspaces, particularly in cross-modal interaction layers. These rotations are associated with reduced task-aligned margins and larger task-aware representation harm under missing-modality inputs. We propose Geodesic Unlearning (GU), a lightweight parameter-editing method that leverages Grassmannian subspace geometry for structured subspace correction to improve missing-modality robustness. It rotates the principal input subspace toward a unimodal reference along a geodesic path. We prove that this correction minimizes the distance to the reference within a fixed subspace-distance budget. Experiments across architectures and datasets show that GU improves performance under missing-modality inference while preserving full-modality accuracy, outperforming strong missing-modality robustness baselines. These findings support a geometric view of deployment-time missing-modality degradation and suggest localized subspace editing as a practical route for robustness correction.
- Abstract(参考訳): モダリティの欠如は、マルチモーダル学習における長年にわたる課題である。
既存の手法は通常、モダリティ回復や適応戦略を通じてこの問題に対処する。
しかし、モデルの内部の相互依存はマルチモーダルトレーニング中に発生し、後に堅牢性を損なう。
完全なモダリティで訓練されたモデルは、推論時に1つのモダリティが欠落しているときに、非モダリティモデルに過小評価することができる。
このパターンは、融合モデル、CLIPスタイルの2towerモデル、ビジョン言語モデルなど、さまざまなアーキテクチャにまたがっている。
このような劣化は、主パラメータ部分空間における学習された相互依存と密接に関連していることを示す。
マルチモーダルトレーニングは、特にクロスモーダル相互作用層において、これらの部分空間の構造的回転を誘導する。
これらの回転は、タスクアライメントのマージンの減少と、モダリティの欠如によるタスクアウェア表現の損失の増大と関連付けられている。
本研究では,GU (Geodesic Unlearning) を提案する。GU (Geodesic Unlearning) は,Grassmann的部分空間幾何を利用して構造的部分空間補正を行い,モダリティの欠落を改善する軽量なパラメータ編集手法である。
主入力部分空間を測地線に沿って一様参照に回転させる。
この補正は、固定された部分空間距離予算内での基準までの距離を最小化する。
アーキテクチャとデータセットをまたいだ実験では、GUは完全なモダリティの正確さを維持しながら、欠落モダリティの推論の下でパフォーマンスを向上し、強力なモダリティのロバストネスベースラインを上回っている。
これらの知見は, 展開時間不足モード劣化の幾何学的視点をサポートし, 局所的な部分空間編集をロバストネス補正の実用的な方法として提案する。
関連論文リスト
- Anisotropic Modality Align [91.23979617826926]
マルチモーダルな大規模言語モデルの訓練は、高品質なペア型マルチモーダルデータの不足により、長い間制限されてきた。
近年の研究では、事前訓練されたマルチモーダルコントラストモデルの共有表現空間がブリッジとして機能し、非モーダルデータを用いたマルチモーダルトレーニングを可能にすることが示されている。
中心となる障害は、共有空間の永続的なモダリティギャップにある。
論文 参考訳(メタデータ) (2026-05-08T14:53:24Z) - Evaluation Before Generation: A Paradigm for Robust Multimodal Sentiment Analysis with Missing Modalities [21.767502810187477]
モダリティの欠如は、マルチモーダルな感情分析において根本的な課題となる。
既存のアプローチは主に、素早い学習と事前訓練されたモデルを通じて堅牢性を改善する。
Promptベースのミスモダリティ適応フレームワークがこれらの問題に対処するために提案されている。
論文 参考訳(メタデータ) (2026-04-07T07:59:06Z) - Guided Verifier: Collaborative Multimodal Reasoning via Dynamic Process Supervision [11.159231524113764]
マルチモーダル大規模言語モデル(MLLM)の複雑な推論能力を高めるための重要なメカニズムとして強化学習(RL)が登場した。
本稿では,これらの構造的制約に対処する textbfGuided Verifier フレームワークを提案する。
我々は,マルチモーダル幻覚をターゲットとした特殊なデータ合成パイプラインを開発し,プロセスレベルの負の textbfCoRe データセットとtextbfCorrect-guide textbfReasoning トラジェクトリを構築し,ガイド付き検証器を訓練する。
論文 参考訳(メタデータ) (2026-02-04T07:38:42Z) - StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval [75.28673512571449]
Continual Text-to-Video Retrievalの重要な課題はフィーチャードリフトだ。
我々はCTVRのための構造化クロスモーダルアライメント手法であるStructAlignを提案する。
我々の手法は、常に最先端の連続検索手法より優れています。
論文 参考訳(メタデータ) (2026-01-28T13:34:44Z) - From Sparse Decisions to Dense Reasoning: A Multi-attribute Trajectory Paradigm for Multimodal Moderation [59.27094165576015]
疎度な意思決定から高密度な推論トレースへ移行する新しい学習パラダイム(UniMod)を提案する。
モノリシックな意思決定タスクを多次元境界学習プロセスに再構成し,エビデンス,モダリティ評価,リスクマッピング,政策決定,応答生成を含む構造化軌道を構築する。
タスク固有のパラメータを分離し、トレーニングダイナミクスを再バランスさせ、マルチタスク学習における多様な目的間の干渉を効果的に解消する、特別な最適化戦略を導入する。
論文 参考訳(メタデータ) (2026-01-28T09:29:40Z) - Switchable Representation Learning Framework with Self-compatibility [50.48336074436792]
自己整合性(SFSC)を考慮した交換可能な表現学習フレームワークを提案する。
SFSCは1つのトレーニングプロセスを通じて、異なる能力を持つ一連の互換性のあるサブモデルを生成する。
SFSCは評価データセット上で最先端のパフォーマンスを達成する。
論文 参考訳(メタデータ) (2022-06-16T16:46:32Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。