論文の概要: Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
- arxiv url: http://arxiv.org/abs/2608.05000v1
- Date: Wed, 05 Aug 2026 16:09:25 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-06 14:48:43.982175
- Title: Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
- Title(参考訳): マルチモーダルプレトレーニングの物理に向けて:知識フロー,モダリティシナジー,早期統一,レシピ
- Authors: Junlin Han, Shengbang Tong, David Fan, Minghao Chen, Philip Torr, Filippos Kokkinos, Mike Lewis,
- Abstract要約: 我々は,マルチモーダル事前学習の体系的な探索を通じて,経験的明瞭度を提供する。
人工および大規模実世界のデータセットに関する我々の実験は、マルチモーダルプレトレーニングの物理に関する4つの重要な洞察をもたらす。
- 参考スコア(独自算出の注目度): 57.55143519213221
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities interact during unified training remain underexplored. We provide empirical clarity through a systematic exploration of multimodal pretraining. Our controlled experiments on both synthetic and large-scale real-world datasets yield four key insights into the physics of multimodal pretraining: (i) Knowledge Flow: We disentangle how language, visual understanding, and visual generation transfer knowledge across modalities, revealing distinct patterns of influence and asymmetry; (ii) Synergy vs. Competition: We show that data "complexity" largely determines whether modalities are synergistic, identify architectural choices that promote synergy: such as shared attention and normalization with modality-specific feed-forward layers, and find that these behaviors generalize across different visual tokenizer designs; (iii) Early Unification: Unifying modalities from the very early stages and training them jointly is shown to be more effective than late alignment or sequential training. This process uncovers a vision laziness phenomenon, where delayed integration leads models to rely on language priors; (iv) Recipes: We derive efficient pretraining recipes that achieve strong generative performance using only 5% of the compute budget. These core findings are subsequently validated at scale by training multiple 13.5B MoE models on 2T tokens. We hope this study provides a principled foundation for understanding and scaling multimodal pretraining.
- Abstract(参考訳): ビジョンは基礎モデルを前進させるために重要な軸を提供し、ネイティブに統一されたマルチモーダル事前訓練へとシフトする。
この運動量にもかかわらず、統一トレーニング中にモダリティがどう相互作用するかという設計空間と基本的なメカニズムは未解明のままである。
我々は,マルチモーダル事前学習の体系的な探索を通じて,経験的明瞭度を提供する。
人工および大規模実世界のデータセットに関する制御実験は、マルチモーダルプレトレーニングの物理に関する4つの重要な洞察をもたらす。
一 知識の流れ 言語、視覚的理解、視覚的生成に関する知識が様々に変化し、異なる影響パターンと非対称性を明らかにすること。
(二)相乗効果対競争:データ「複雑性」は、相乗効果が相乗効果であるか否かを判断し、相乗効果を促進するアーキテクチャ的選択を同定する。
(三)初期統一:初期からモダリティを統一し、それらを共同で訓練することは、後期調整や連続訓練よりも効果的であることが示されている。
このプロセスは、遅延統合によってモデルが言語優先に依存するという視覚遅延現象を明らかにする。
(4)レシピ:計算予算の5%しか使わず、高い生成性能を達成する効率的な事前学習レシピを導出する。
これらのコアはその後、2Tトークン上で複数の13.5B MoEモデルをトレーニングすることで大規模に検証される。
この研究がマルチモーダル事前学習の理解とスケーリングの基盤となることを願っている。
関連論文リスト
- Beyond Language Modeling: An Exploration of Multimodal Pretraining [125.34714978184638]
我々は、制御されたオフスクラッチ事前学習実験を通して経験的明瞭度を提供する。
我々はトランスフュージョン・フレームワークを採用し、言語と視覚の拡散を次々に予測する。
我々は、MoEアーキテクチャが、言語によって要求される高いモデル容量を提供することにより、このスケーリング非対称性を調和させることを実証する。
論文 参考訳(メタデータ) (2026-03-03T18:58:00Z) - MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings [75.0617088717528]
MoCaは、トレーニング済みのVLMバックボーンを効果的な双方向埋め込みモデルに変換するためのフレームワークである。
MoCaは、MMEBとViDoRe-v2ベンチマークのパフォーマンスを継続的に改善し、新しい最先端の結果を達成する。
論文 参考訳(メタデータ) (2025-06-29T06:41:00Z) - Will Pre-Training Ever End? A First Step Toward Next-Generation Foundation MLLMs via Self-Improving Systematic Cognition [89.50068130832635]
自己改善認知 (SIcog) は、マルチモーダル知識によって次世代のMLLMを構築するための自己学習フレームワークである。
ステップバイステップの視覚的理解のためのChain-of-Descriptionを提案し、詳細なマルチモーダル推論をサポートするために構造化されたChain-of-Thought(CoT)推論を統合する。
実験は、マルチモーダル認知を増強したMLLMの開発におけるSIcogの有効性を示す。
論文 参考訳(メタデータ) (2025-03-16T00:25:13Z) - A Statistical Theory of Contrastive Pre-training and Multimodal Generative AI [35.067337756433965]
マルチモーダル生成AIシステムは、異なるモーダルの表現を学ぶために、対照的な事前学習に依存している。
本稿では、下流タスクにおける対照的な事前学習の成功を説明するための理論的枠組みを開発する。
論文 参考訳(メタデータ) (2025-01-08T17:47:06Z) - Diving into Self-Evolving Training for Multimodal Reasoning [36.70979791148913]
自己進化的トレインは複雑な推論タスクの鍵となるアプローチとして登場した。
本稿では,強化学習のレンズによるマルチモーダル推論のための自己進化学習を再構成する。
M-STARは、様々なサイズと多様なベンチマークのモデル間で一貫したパフォーマンス向上を実現するフレームワークである。
論文 参考訳(メタデータ) (2024-12-23T10:18:41Z) - i-Code: An Integrative and Composable Multimodal Learning Framework [99.56065789066027]
i-Codeは、視覚、音声、言語を統一的で汎用的なベクトル表現に柔軟に組み合わせられる自己教師型事前学習フレームワークである。
システム全体は、マスク付きモダリティ・ユニット・モデリングやクロスモダリティ・コントラスト・ラーニングなどの新しい目的により、エンドツーエンドで事前訓練されている。
実験の結果、i-Codeは5つのビデオ理解タスクとGLUE NLPベンチマークで最先端技術を上回る性能を示し、最大11%改善した。
論文 参考訳(メタデータ) (2022-05-03T23:38:50Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。