論文の概要: DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models
- arxiv url: http://arxiv.org/abs/2607.08434v1
- Date: Thu, 09 Jul 2026 12:54:05 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-10 14:45:27.539906
- Title: DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models
- Title(参考訳): DeltaV: 統合された大規模マルチモーダルモデルにおけるビジュアル状態のアップデートを考える
- Authors: Pengjie Wang, Linger Deng, Zujia Zhang, Shaojie Zhang, Zhenbo Luo, Pei Fu, Jian Luan, Xiang Bai, Yuliang Liu,
- Abstract要約: 現在のUnified Large Multimodal Models (ULMM) は、テキスト推論と中間視覚状態によるインターリーブされたマルチモーダル推論をサポートする。
このフルイメージ生成パラダイムは、実質的な視覚的冗長性を導入し、スパースで推論クリティカルな状態遷移を監督する。
フルイメージ生成を視覚的更新に置き換えるULMMであるDeltaVを提案する。
- 参考スコア(独自算出の注目度): 61.3037243251433
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Current Unified Large Multimodal Models (ULMMs) support interleaved multimodal reasoning through textual reasoning and intermediate visual states, but typically generate each visual state as a full image. This full-image generation paradigm introduces substantial visual-token redundancy and dilutes supervision on sparse yet reasoning-critical state transitions. We propose DeltaV, a ULMM that replaces full-image generation with visual updates. Conditioned on historical visual states, DeltaV incrementally predicts compact update tokens that capture the visual changes across reasoning steps, avoiding repeated modeling of unchanged content. To align the token budget of each update with the magnitude of visual change, DeltaV introduces a temporal similarity (TSIM) Router, which stops allocating tokens once the marginal reconstruction gain falls below a threshold. To support more diverse and generalizable reasoning, we further construct StructCoT, a large-scale interleaved multimodal reasoning dataset with 1.05M samples spanning 44 task domains. Experiments show that the visual-update paradigm reduces newly generated visual tokens by 55.6\% on average without compromising reconstruction fidelity, and improves multimodal reasoning by 3.3\% over full-image generation. Trained with StructCoT and large-scale multimodal data, DeltaV-2B further outperforms substantially larger open-source models by 8.4\% on in-domain multimodal reasoning evaluations and surpasses the comparable-scale Qwen3-VL-2B by 5.9\% on external multimodal reasoning and understanding benchmarks. Code, models, and StructCoT will be released at https://github.com/Pengjie-W/DeltaV.
- Abstract(参考訳): 現在のUnified Large Multimodal Model(ULMM)は、テキスト推論と中間的な視覚状態を通じてインターリーブされたマルチモーダル推論をサポートするが、通常、各視覚状態をフルイメージとして生成する。
このフルイメージ生成パラダイムは、実質的な視覚的冗長性を導入し、スパースで推論クリティカルな状態遷移を監督する。
フルイメージ生成を視覚的更新に置き換えるULMMであるDeltaVを提案する。
DeltaVは、歴史的視覚状態に基づいて、推論ステップ間で視覚的変化をキャプチャするコンパクトな更新トークンを段階的に予測する。
各更新のトークン予算を視覚的変化の大きさに合わせるため、DeltaVは時間的類似性(TSIM)ルータを導入する。
より多種多様な一般化可能な推論をサポートするために、44のタスクドメインにまたがる1.05Mサンプルを用いた大規模インターリーブマルチモーダル推論データセットであるStructCoTを構築した。
実験により、視覚更新パラダイムは、再建忠実度を損なうことなく、新たに生成された視覚トークンを平均55.6%削減し、フル画像生成よりも3.3倍のマルチモーダル推論を改善することが示された。
StructCoTと大規模マルチモーダルデータを用いて訓練されたDeltaV-2Bは、ドメイン内のマルチモーダル推論評価では8.4 %、外部マルチモーダル推論と理解ベンチマークでは5.9 %の精度でQwen3-VL-2Bを上回った。
コード、モデル、StructCoTはhttps://github.com/Pengjie-W/DeltaV.comでリリースされる。
関連論文リスト
- Revisiting Change VQA in Remote Sensing with Structured and Native Multimodal Qwen Models [2.4564286418294468]
視覚的質問応答の変化(Change VQA)は、バイテンポラルリモートセンシング(RS)画像間の意味的変化に関する自然言語質問に答える問題に対処する。
近年の視覚言語モデル(VLM)は、時間的RS画像理解のために研究されているが、現代マルチモーダルモデルの文脈において、変化VQAは未解明のままである。
Qwen3-VLは、多次元視覚条件とフルアテンションデコーダを備えた構造化視覚言語パイプラインと、単一ステージアライメントとハイブリッドデコーダバックボーンを組み合わせたネイティブマルチモーダルモデルであるQwen3.5を比較した。
論文 参考訳(メタデータ) (2026-04-20T15:47:52Z) - Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation [66.53544128707817]
Cheersは、パッチレベルの詳細をセマンティック表現から切り離す、統一されたマルチモーダルモデルである。
チェアは視覚的理解と生成の両方において、高度なUMMと一致または超えます。
論文 参考訳(メタデータ) (2026-03-13T08:55:27Z) - Kelix Technical Report [86.64551727600104]
我々は、完全離散自己回帰統一モデルであるKelixを紹介し、離散的および連続的な視覚表現間の理解ギャップを埋める。
最近の研究は、完全自己回帰型マルチモーダルモデリングを可能にするために、離散的な視覚的トークン化を探求している。
論文 参考訳(メタデータ) (2026-02-10T14:48:26Z) - Multi-Modal View Enhanced Large Vision Models for Long-Term Time Series Forecasting [53.332533610841885]
時系列は画像やテキストに変換でき、同じ信号のマルチモーダルビュー(MMV)を提供する。
これらのMMVは相補的なパターンを明らかにし、長期時系列予測(LTSF)のための大型ビジョンモデル(LVM)のような強力な事前訓練された大規模モデルの使用を可能にする。
DMMVは、トレンド・シーズンの分解と新しいバックキャスト・レジデンシャル・アダプティブ・コンダプティブ・コンダプションを活用し、LTSFのためのMMVを統合する新しい分解ベースマルチモーダル・ビュー・フレームワークである。
論文 参考訳(メタデータ) (2025-05-29T20:55:24Z) - v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning [27.688428439248607]
簡単なポイント・アンド・コピーアプローチによるアクティブな視覚的参照を可能にする軽量な拡張であるv1を紹介する。
これにより、モデルは関連するイメージパッチを特定し、埋め込みを推論ストリームにコピーすることができる。
我々のポインティング戦略では、MLLMはセマンティックな表現をキーとして直接イメージパッチを選択でき、知覚的証拠はモデルの推論と同じ空間に埋め込まれている。
論文 参考訳(メタデータ) (2025-05-24T19:30:47Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。