論文の概要: Recursive Self-Improvement in Unified Multimodal Models
- arxiv url: http://arxiv.org/abs/2610.03002v1
- Date: Fri, 02 Oct 2026 08:34:21 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-06 00:14:30.285961
- Title: Recursive Self-Improvement in Unified Multimodal Models
- Title(参考訳): 統一マルチモーダルモデルにおける再帰的自己改善
- Abstract要約: 統一マルチモーダルモデル(UMM)は、テキストと画像の両方を理解し、生成する。
UMMのテキストと視覚能力が相互にトレーニングデータを供給する訓練ループを提案する。
- 参考スコア(独自算出の注目度): 41.20184334137855
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Unified multimodal models (UMMs) understand and generate both text and images, which lets a model produce its own training data. Existing self-improvement in UMMs keeps supervision on the visual side, where image understanding judges image generation. We propose recursive cross-capability self-improvement (RSI), a training loop in which the text and visual abilities of a UMM supply training data for one another. In each round, the model generates images and reads them to find where it falls short. It then writes programs aimed at these shortcomings, and execution verifies every result against its specification. Verified renders train image generation, while labeled renders and the model's own correct programs train visual understanding and program writing. Program execution thus acts as a source of truth outside the model, so errors do not accumulate across rounds. We study RSI on charts and build BasicChartBench to evaluate open models early in training. On requests worded differently from training, four rounds of RSI raise the score from 45.7% to 60.2%, while continued training stays at 46.3%. Verified construction carries most of the gain, and targeting the model's failures adds 3.5%. Along the way, the share of verified programs rises from 48.9% to 95.2%, and the reader's accuracy on edited renders rises from 55.6% to 87.4%.
- Abstract(参考訳): 統一マルチモーダルモデル(UMM)は、テキストと画像の両方を理解し、生成する。
UMMにおける既存の自己改善は、画像理解が画像生成を判断する視覚面を監督し続ける。
本稿では,UMMのテキストと視覚能力を相互に供給する訓練ループである,再帰的クロスキャパビリティ自己改善(RSI)を提案する。
各ラウンドで、モデルが画像を生成し、それを読み取り、それが不足している場所を見つける。
そして、これらの欠点を目的としたプログラムを書き、実行はその仕様に対してすべての結果を検証する。
検証されたレンダリングはトレーニングイメージの生成を、ラベル付きレンダリングはモデル自身の正しいプログラムは視覚的理解とプログラム記述を訓練する。
したがって、プログラムの実行はモデル外の真理の源として機能するので、エラーはラウンド全体で蓄積されない。
我々は、チャート上でRSIを研究し、BasicChartBenchを構築し、訓練の早い段階でオープンモデルを評価する。
RSIの4ラウンドは45.7%から60.2%に上昇し、継続トレーニングは46.3%にとどまった。
検証された建設は、ほとんどの利益をもたらし、モデルの失敗を狙うと3.5%も増加します。
その過程で、検証されたプログラムのシェアは48.9%から95.2%に上昇し、編集されたレンダリングの正確さは55.6%から87.4%に上昇した。
関連論文リスト
- Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs [23.958966900531692]
MLLM(Multimodal large language model)は、画像として表示されるテキストを処理できるが、同じコンテンツがテキストトークンとして提供される場合よりも処理が悪くなることが多い。
我々は,この「モダリティギャップ」を7つのベンチマークを5つの入力モードで評価することにより,系統的に診断する。
論文 参考訳(メタデータ) (2026-03-10T02:14:23Z) - ViUniT: Visual Unit Tests for More Robust Visual Programming [104.55763189099125]
モデルが正しく答えると、不正なプログラムを33%生成します。
自動単体テストを生成することで、視覚プログラムの信頼性を向上させるためのフレームワークであるVisual Unit Testing (ViUniT)を提案する。
論文 参考訳(メタデータ) (2024-12-12T01:36:18Z) - ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models [103.25208095165486]
既存のプラクティスは命令データを生成するために、強力だが高価な言語モデル(LLM)やマルチモーダル言語モデル(MLM)に依存している。
本稿では,シーングラフを画像のシンボル表現として利用し,視覚中心の命令データを体系的に合成するプログラムを提案する。
提案手法は,データ生成プロセスの解釈可能性と制御性を保証し,実際の精度を維持しながら効率よくスケールする。
論文 参考訳(メタデータ) (2024-12-09T21:44:02Z) - Composing Ensembles of Pre-trained Models via Iterative Consensus [95.10641301155232]
本稿では,異なる事前学習モデルのアンサンブルを構成するための統一的なフレームワークを提案する。
事前学習したモデルを「ジェネレータ」あるいは「スコーラ」として使用し、クローズドループ反復コンセンサス最適化により構成する。
スコアラーのアンサンブルによって達成されたコンセンサスは、シングルスコアラーのフィードバックよりも優れていることを示す。
論文 参考訳(メタデータ) (2022-10-20T18:46:31Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。