論文の概要: Generalize LMMs to Versatile Visual Modalities via Fabricated Modality Synthesis
- arxiv url: http://arxiv.org/abs/2607.10308v1
- Date: Sat, 11 Jul 2026 13:28:16 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-14 15:40:48.390413
- Title: Generalize LMMs to Versatile Visual Modalities via Fabricated Modality Synthesis
- Title(参考訳): ファブリケートモダリティ合成による多目的視覚モダリティへのLMMの一般化
- Authors: Shihao Yuan, Yuanze Li, Ruyi Zhang, Ming Liu, Wangmeng Zuo,
- Abstract要約: モーダリティ合成とモーダリティコンテキストを通してLMMに能力を持たせるためのトレーニングフレームワーク VVM-Tuning を提案する。
我々は、RGBシーンから多彩な外見変化画像を合成し、異なる視覚的外観から不変のセマンティクスをアンタングルするためにモデルを訓練する。
次に、アクセプションにモダリティコンテキストを導入し、モデルがこれらの外観変化をモダリティ関連属性にマッピングするのを支援するために、インストラクションチューニングを使用する。
- 参考スコア(独自算出の注目度): 45.48024663735415
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Despite the advancements of Large Multimodal Models (LMMs) in RGB vision, their ability to generalize to unseen visual modalities remains a largely unexplored challenge. We argue that different visual modalities are merely distinct samplings of the same physical world. Therefore, effective generalization requires models to possess both modality-agnostic perception of scene semantics and the adaptability to modality-specific characteristics. To achieve this, we propose a training framework, VVM-Tuning, to equip LMMs with these capabilities through modality synthesis and modality contexts. Specifically, we synthesize diverse appearance-varied images from RGB scenes, training the model to disentangle invariant semantics from varying visual appearances, and align these appearances with language for visual concepts decoupled from modalities. We then introduce modality contexts in the prompt and use instruction tuning to assist the model in mapping these appearance variations back to modality-related attributes, enabling zero-shot adaptation to unseen modalities during inference. To facilitate research in this direction, we introduce VVM-Bench, a comprehensive benchmark featuring 6 real and synthetic modalities to evaluate semantic perception and modality understanding. Experiments demonstrate that, via our training on synthetic modalities, 5 tested models exhibit consistent improvements on both real-world and novel synthetic modalities without in-modality training. Source code and data will be publicly available at https://github.com/Hunter-Will/VVM-Tuning.
- Abstract(参考訳): RGBビジョンにおけるLMM(Large Multimodal Models)の進歩にもかかわらず、見えない視覚的モダリティに一般化する能力は、まだ明らかにされていない課題である。
異なる視覚的モダリティは、単に同じ物理世界の異なるサンプリングである、と我々は主張する。
したがって、効果的な一般化には、シーン意味論のモダリティ非依存的な認識と、モダリティ固有の特徴への適応性の両方を持つモデルが必要である。
そこで本研究では,モダリティ合成とモダリティコンテキストを通じて,LMMにこれらの機能を持たせるためのトレーニングフレームワーク VVM-Tuning を提案する。
具体的には、RGBシーンから多彩な外観変化画像を合成し、様々な視覚的外観から不変意味論を切り離すようモデルを訓練し、これらの外観をモダリティから切り離された視覚概念のための言語と整合させる。
次に、インプロンプトにモダリティコンテキストを導入し、モデルがこれらの外見のバリエーションをモダリティ関連属性にマッピングするのを支援するために、インストラクション中に目に見えないモダリティへのゼロショット適応を可能にするために、インストラクションチューニングを使用する。
この方向の研究を容易にするために,6つの実・合成モダリティを特徴とする総合的なベンチマークであるVVM-Benchを導入し,意味認識とモダリティ理解を評価する。
実験により, 実世界および新奇な合成モダリティの両面において, 模擬モダリティのトレーニングにより, 実世界および新奇な合成モダリティに一貫した改善が認められた。
ソースコードとデータはhttps://github.com/Hunter-Will/VVM-Tuning.comで公開されている。
関連論文リスト
- Towards Understanding Multimodal Fine-Tuning: Spatial Features [25.349396112139214]
Vision-Language Models (VLM) は、事前訓練された言語モデルとビジョンエンコーダをペアリングすることで、幅広いタスクにおいて強力なパフォーマンスを達成する。
本稿では,ステージワイドモデル差分法によるVLM適応の最初の力学解析について述べる。
論文 参考訳(メタデータ) (2026-02-06T18:48:18Z) - MoRA: Missing Modality Low-Rank Adaptation for Visual Recognition [22.42423566934899]
MoRAはパラメータ効率のよい微調整法であり、クロスモーダル相互作用を明示的にモデル化する。
MoRAは、欠落したモダリティシナリオにおける平均的なパフォーマンス改善を5.24%向上させる。
論文 参考訳(メタデータ) (2025-11-09T04:52:42Z) - Unified modality separation: A vision-language framework for unsupervised domain adaptation [60.8391821117794]
教師なしドメイン適応(Unsupervised domain adapt, UDA)は、ラベル付きソースドメインでトレーニングされたモデルが新しいラベル付きドメインを扱うことを可能にする。
本稿では,モダリティ固有成分とモダリティ不変成分の両方に対応可能な統一モダリティ分離フレームワークを提案する。
提案手法は,9倍の計算効率で最大9%の性能向上を実現している。
論文 参考訳(メタデータ) (2025-08-07T02:51:10Z) - Analyzing Finetuning Representation Shift for Multimodal LLMs Steering [56.710375516257876]
隠れた状態を解釈可能な視覚的概念とテキスト的概念にマッピングすることを提案する。
これにより、オリジナルモデルや微調整モデルからのシフトなど、特定のセマンティックダイナミクスをより効率的に比較することが可能になります。
また,これらの変化を捉えるためにシフトベクトルを用いることを実証する。
論文 参考訳(メタデータ) (2025-01-06T13:37:13Z) - LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering [30.51487692912812]
MLLM(Multimodal Large Language Models)は、大規模言語モデル(LLM)に視覚表現を統合することで、視覚的タスクを大幅に進歩させる。
目的を達成するためにモダリティリニア表現ステアリング(MoReS)を導入する。
MoReSはモデル全体の固有のモダリティを効果的に再バランスさせ、そこでキーとなるアイデアは、各モデル層をまたいだ視覚部分空間の線形変換を通じて視覚表現を操ることである。
論文 参考訳(メタデータ) (2024-12-16T21:14:11Z) - OmniBench: Towards The Future of Universal Omni-Language Models [63.16606414452612]
OmniBenchは、視覚的、音響的、テキスト的入力を同時に認識し、解釈し、推論する能力を評価するために設計された新しいベンチマークである。
評価の結果,オープンソース OLM は三モーダル文脈における命令追従や推論に重大な制限があることが明らかとなった。
我々は,OLM性能を向上させるため,より堅牢な3モーダル統合技術とトレーニング戦略の開発を提唱する。
論文 参考訳(メタデータ) (2024-09-23T17:59:05Z) - Learning Modality-agnostic Representation for Semantic Segmentation from Any Modalities [8.517830626176641]
Any2Segは、任意の視覚的条件におけるモダリティの組み合わせから堅牢なセグメンテーションを実現する新しいフレームワークである。
4つのモダリティを持つ2つのベンチマークの実験は、Any2Segがマルチモーダル設定の下で最先端を達成することを示した。
論文 参考訳(メタデータ) (2024-07-16T03:34:38Z) - Exploiting modality-invariant feature for robust multimodal emotion
recognition with missing modalities [76.08541852988536]
我々は、欠落したモダリティ・イマジネーション・ネットワーク(IF-MMIN)に不変な特徴を用いることを提案する。
提案モデルは,不確実なモダリティ条件下で,すべてのベースラインを上回り,全体の感情認識性能を不変に向上することを示す。
論文 参考訳(メタデータ) (2022-10-27T12:16:25Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。