論文の概要: Mixture of Probes: Learning from Privileged Modalities in Multimodal LLMs Through Probing
- arxiv url: http://arxiv.org/abs/2607.08839v1
- Date: Thu, 09 Jul 2026 18:00:47 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-13 14:47:12.722552
- Title: Mixture of Probes: Learning from Privileged Modalities in Multimodal LLMs Through Probing
- Title(参考訳): プローブの混合:プローブによるマルチモーダルLLMの原始モードからの学習
- Abstract要約: MoP(Mixture of Probes)は、MLLM(Multimodal Large Language Models)内のモダリティ固有信号とモダリティ一般信号を切り離す新しいフレームワークである。
MoPはMLLMベースラインを一貫して上回り、最大65%の改善を実現している。
- 参考スコア(独自算出の注目度): 18.589015768150578
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Multimodal Large Language Models (MLLMs) are typically designed under the assumption that all modalities available during training will also be accessible at inference. However, many real-world settings violate this assumption, requiring models to operate under a privileged modality setting, where auxiliary modalities are available only during training. While these modalities contain valuable information, existing MLLMs largely fail to leverage them effectively, as they treat modalities as interchangeable inputs rather than sources of complementary supervision. We propose Mixture of Probes (MoP), a novel framework that disentangles modality-specific and modality-general signals within the MLLM, allowing the model to preserve modality-dependent structure while learning transferable representations across modalities. At its core, MoP achieves this through a structured probing mechanism that extracts and organizes information from intermediate representations of a shared modality encoder, rather than relying only on final-layer alignment as done in existing MLLMs. To support this disentanglement, we further introduce MoP Cross-modal Training (MoP-X), a training strategy for MoP centered around a probe disentanglement loss that prevents probe collapse and encourages cross-modal learning. We evaluate MoP across two domains spanning eight tasks and four modalities under a comprehensive evaluation protocol tailored to the privileged modality setting, where each modality is independently treated as the sole input at inference time. MoP consistently outperforms strong MLLM baselines, achieving up to 65% relative improvement, demonstrating that auxiliary modalities, even when unavailable at inference, can provide substantial gains when effectively leveraged during training. Code, model checkpoints, and evaluation protocols will be made available at https://github.com/Sony/MoP.
- Abstract(参考訳): MLLM(Multimodal Large Language Models)は通常、トレーニング中に利用できるすべてのモダリティが推論時にもアクセス可能であるという前提の下で設計される。
しかし、多くの現実世界の設定はこの前提に反し、訓練中にのみ補助的なモダリティが利用できる特権的なモダリティ設定の下でモデルを操作する必要がある。
これらのモダリティには貴重な情報が含まれているが、既存のMLLMはそれらを効果的に活用することができない。
MLLM内のモダリティ固有信号とモダリティ一般信号とをアンタングル化する新しいフレームワークであるMixture of Probes(MoP)を提案し,モダリティ間の伝達可能な表現を学習しながら、モデルがモダリティ依存構造を維持できるようにする。
コアとなるMoPは、既存のMLLMのように最終層アライメントのみに依存するのではなく、共有モダリティエンコーダの中間表現から情報を抽出し、整理する構造化されたプローブ機構によってこれを達成している。
そこで本研究では,MoPクロスモーダルトレーニング(MoP-X)についても紹介する。これは,プローブの崩壊を防止し,クロスモーダル学習を促進するプローブディアングルメント損失を中心としたMoPトレーニング戦略である。
我々は,8つのタスクにまたがる2つの領域にまたがるMoPを評価し,各モードを推論時に単独の入力として独立に扱う特権的モダリティ設定に適合する包括的評価プロトコルで4つのモダリティを評価する。
MoPは強力なMLLMベースラインを一貫して上回り、最大65%の相対的な改善を達成し、推論で利用できない場合でも補助的なモダリティがトレーニング中に効果的に活用された場合、かなりの利益をもたらすことを示した。
コード、モデルチェックポイント、評価プロトコルはhttps://github.com/Sony/MoP.comで利用可能になる。
関連論文リスト
- Multimodal LLMs under Pairwise Modalities [16.75545711899814]
ペアワイズデータのみを用いて、モダリティ間で潜在表現を整列する表現学習フレームワークを提案する。
特に第1段階では、自己モダル再構成とペアワイドコントラスト学習の両方により、モダリティ間の共有潜在空間を学習する。
ステージ2では、新たに導入されたモダリティのエンコーダと事前訓練されたモダリティのデコーダを統合し、クロスモーダル転送と生成を容易にする。
論文 参考訳(メタデータ) (2026-05-20T11:44:01Z) - LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete Multimodal Observations [53.20772659095155]
本稿では、トレーニング時不完全観察において、より困難なIMLの設定に取り組む。
本稿では,この課題を条件付きシーケンス推論タスクとして再構成したLIMSSR(LLM-Driven Incomplete Multimodal Sequence-to-Score Reasoning)を提案する。
論文 参考訳(メタデータ) (2026-05-01T06:11:42Z) - MDL: A Unified Multi-Distribution Learner in Large-scale Industrial Recommendation through Tokenization [14.534152704620261]
産業レコメンデータシステムは、多様なユーザインタラクションやコンテキストを扱うために、MSL(Multi-scenario Learning)とMulti-task Learning(MTL)を採用するようになっている。
既存のアプローチでは,(1)複雑な特徴モジュールとの相互作用が限られているため,大規模モデルパラメータの非活用,(2)統合されたフレームワークにおけるシナリオとタスク情報の共同モデリングの難しさ,という2つの重大な欠点がある。
大規模言語モデル(LLM)における「プロンプト」パラダイムにインスパイアされた、統一された textbfMulti-textbfDistribution textbfL MSL フレームワークを提案する。
論文 参考訳(メタデータ) (2026-02-07T12:34:27Z) - Sample-efficient Integration of New Modalities into Large Language Models [48.81776019848246]
マルチモーダル基礎モデルはいくつかのモダリティを処理できる。
本稿では,大規模言語モデルへのサンプル効率改善手法を提案する。
SEMIは、新しいモダリティを数秒で統合することで、サンプル効率を大幅に向上することがわかった。
論文 参考訳(メタデータ) (2025-09-04T18:41:59Z) - Dynamic Modality Scheduling for Multimodal Large Models via Confidence, Uncertainty, and Semantic Consistency [0.0]
本研究では,各モダリティの寄与をサンプル単位で適応的に調整する新しいフレームワークである動的モダリティスケジューリング(DMS)を提案する。
VQA、画像テキスト検索、キャプションタスクの実験結果から、DMSはクリーンとロバストの両方のパフォーマンスを著しく改善することが示された。
論文 参考訳(メタデータ) (2025-06-15T05:15:52Z) - OptMerge: Unifying Multimodal LLM Capabilities and Modalities via Model Merging [124.91183814854126]
モデルマージは、複数のエキスパートモデルをひとつのモデルに組み合わせようとしている。
本稿ではMLLMのトレーニングと評価のタスクを明確に分割したモデルマージ研究のベンチマークを紹介する。
モデルマージは、トレーニングデータを必要とせずに改善されたMLLMを構築するための有望な方法であることがわかった。
論文 参考訳(メタデータ) (2025-05-26T12:23:14Z) - LLMs Can Evolve Continually on Modality for X-Modal Reasoning [62.2874638875554]
既存の手法は、モーダル固有の事前訓練とジョイント・モーダルチューニングに大きく依存しており、新しいモーダルへと拡張する際の計算上の負担が大きくなった。
PathWeaveは、Modal-Path sWitchingとExpAnsion機能を備えた柔軟でスケーラブルなフレームワークである。
PathWeaveは最先端のMLLMと互換性があり、パラメータトレーニングの負担を98.73%削減する。
論文 参考訳(メタデータ) (2024-10-26T13:19:57Z) - On-the-fly Modulation for Balanced Multimodal Learning [53.616094855778954]
マルチモーダル学習は、異なるモーダルからの情報を統合することでモデル性能を向上させることが期待されている。
広く使われている共同トレーニング戦略は、不均衡で最適化されていないユニモーダル表現につながる。
そこで本研究では,OGM(On-the-fly Prediction Modulation)とOGM(On-the-fly Gradient Modulation)の戦略を提案する。
論文 参考訳(メタデータ) (2024-10-15T13:15:50Z) - Multimodal Federated Learning with Missing Modality via Prototype Mask
and Contrast [23.936677199734213]
本稿では,FedAvgベースのFederated Learningフレームワークにプロトタイプライブラリを導入する。
提案手法は,タスク校正されたトレーニング損失とモデルに依存しない一様性推論戦略を定式化するために,欠落したモダリティを表すマスクとしてプロトタイプを利用する。
ベースラインと比較して,トレーニング中に50%のモダリティが欠落し,一様性推論時に23.8%の精度で推論精度が3.7%向上した。
論文 参考訳(メタデータ) (2023-12-21T00:55:12Z) - Unified Multi-modal Unsupervised Representation Learning for
Skeleton-based Action Understanding [62.70450216120704]
教師なしの事前訓練は骨格に基づく行動理解において大きな成功を収めた。
我々はUmURLと呼ばれる統一マルチモーダル非教師なし表現学習フレームワークを提案する。
UmURLは効率的な早期融合戦略を利用して、マルチモーダル機能を単一ストリームで共同でエンコードする。
論文 参考訳(メタデータ) (2023-11-06T13:56:57Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。