論文の概要: Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models
- arxiv url: http://arxiv.org/abs/2608.26317v1
- Date: Wed, 26 Aug 2026 18:48:52 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-28 16:30:58.157256
- Title: Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models
- Title(参考訳): Modality Maturity Index:Omniモデルのマルチモーダル能力を評価するためのベンチマーク
- Abstract要約: 本稿では,MMI(Modality Maturity Index)を提案する。
MMIは853の質問から成り、それぞれが複数の入力モダリティの理解を示すためにモデルを必要とするように慎重に設計されている。
- 参考スコア(独自算出の注目度): 5.886031439164172
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Frontier language models are increasingly marketed as omni systems that can perceive and respond across modalities. Existing evaluation frameworks, however, focus almost exclusively on bimodal understanding, typically text plus one other modality. We propose the Modality Maturity Index (MMI), a benchmark designed to evaluate the multimodal capabilities of large language models across five modalities (text, image, audio, video and document) and combinations of up to three modalities in both inputs and outputs. MMI consists of 893 questions, each carefully crafted to require the model to demonstrate its understanding of multiple input modalities and to generate responses that incorporate various output formats. The questions are designed to be self-contained, with clear expectations for the correct modality or mix of modalities required for an accurate response. Every MMI prompt carries human-authored rubric criteria for each output modality expected in the response; a model's MMI Value expresses the average of the per-modality scores for each prompt. Because low scores can reflect either failure to generate a modality (lack of presence) or failure to generate correct content, we introduce also a supplementary Modality Presence Score (MPS), a per-prompt F1 over the expected output modalities. Applying MMI to five frontier multimodal models, we find that the MPS ranges from only 15.6 (Claude Opus 4.6) to 34.9 (GPT-5.4). Given the low availability of returned modalities to even grade, we report MPS as our main result pending model improvements. To assess the viability of judging output correctness with LLM judges and rubrics, we run a separate experiment with custom generation tools. On the assets that generates, we find that an LLM judge applying the rubrics agrees with rubric-blind human annotators (who score the outputs directly and never see the criteria) on 70.8% of judgments.
- Abstract(参考訳): 最前線の言語モデルは、モダリティにまたがる知覚と応答が可能なオムニシステムとして、ますます市場に出回っている。
しかし、既存の評価フレームワークは、ほとんどバイモーダルな理解(典型的にはテキストと他のモダリティ)に焦点を当てている。
MMI(Modality Maturity Index)は,5つのモダリティ(テキスト,画像,音声,ビデオ,文書)と最大3つのモダリティの組み合わせを対象とする,大規模言語モデルのマルチモーダル能力を評価するためのベンチマークである。
MMIは853の質問で構成され、各質問は、複数の入力モダリティの理解と、様々な出力形式を含む応答を生成するために、モデルを必要とするように慎重に設計されている。
質問は自己完結するように設計されており、正確な応答に必要なモダリティやモダリティの混合に対する明確な期待がある。
すべてのMMIプロンプトは、応答で期待される各出力モダリティに対する人間によるルーリック基準を持ち、モデルのMMI値は各プロンプトに対するモダリティスコアの平均を表す。
低スコアは、モダリティ生成の失敗(プレゼンス不足)や正しいコンテンツ生成の失敗を反映できるため、期待される出力モダリティに対するプロンプトF1である補足的なModality Presence Score(MPS)も導入する。
MMI を5つのフロンティアマルチモーダルモデルに適用すると、MPS は 15.6 (Claude Opus 4.6) から 34.9 (GPT-5.4) までの範囲であることがわかった。
返却モダリティの低レベル化を考えると,MPSはモデルの改良が待たれている主な結果である。
LLM判定器と潤滑剤を用いて出力正当性を判定する可能性を評価するため、我々はカスタム生成ツールを用いて別々に実験を行った。
生成する資産について、LLM判事がルーブリックを施すと、70.8%の判定で、ルーブリックブラインドヒトアノテータ(アウトプットを直接得点し、基準を決して見ていない)と一致していることが分かる。
関連論文リスト
- MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models [70.34265674686516]
マルチモーダル埋め込みモデルは、テキスト、画像、ビデオ、オーディオなどの異種入力を共有意味空間にマッピングすることを目的としている。
本稿では,テキスト,画像,ビデオ,オーディオ,エージェント中心のシナリオにまたがる埋め込みを評価するベンチマークであるMMEB-V3を紹介する。
本研究は, 完全モダリティ埋め込みの系統的解析を行い, 3つの重要な知見を同定する。
論文 参考訳(メタデータ) (2026-04-25T14:15:05Z) - UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models [22.508414355245275]
我々は,新しい,高品質で統一されたオムニモデルベンチマーク,UNO-Benchを紹介する。
このベンチマークは、統一された能力分類の下で、UNi-modalとOmni-modalの両方の能力を効果的に評価するために設計されている。
1250人のオムニモダルの培養サンプルと98%のクロスモーダル可溶性、2480の強化されたユニモーダルサンプルを含んでいる。
論文 参考訳(メタデータ) (2025-10-21T06:14:40Z) - MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation [81.26818054877658]
MMMGは、4つのモダリティの組み合わせにまたがるマルチモーダル生成の包括的なベンチマークである。
人間の評価と高度に一致し、平均94.3%の合意を達成している。
GPTイメージは画像生成の精度は78.3%であるが、マルチモーダル推論とインターリーブ生成では不足している。
論文 参考訳(メタデータ) (2025-05-23T08:21:28Z) - Benchmarking Multi-modal Semantic Segmentation under Sensor Failures: Missing and Noisy Modality Robustness [61.87055159919641]
マルチモーダルセマンティックセグメンテーション(MMSS)は、モーダル間で補完情報を統合することで、単一モーダルデータの制限に対処する。
顕著な進歩にもかかわらず、マルチモーダルデータ品質の変動と不確実性により、研究と実世界の展開の間に大きなギャップが持続する。
Intire-Missing Modality (EMM)、Random-Missing Modality (RMM)、Noisy Modality (NM)の3つのシナリオでMMSSモデルを評価する頑健性ベンチマークを導入する。
論文 参考訳(メタデータ) (2025-03-24T08:46:52Z) - Judge Anything: MLLM as a Judge Across Any Modality [43.51517213949702]
本稿では,タスクAnything と JudgeAnything という2つのベンチマークを導入し,MLLM の全体性能と判断能力を評価する。
TaskAnythingは15のあらゆるモダリティカテゴリでMMUとMMGの機能を評価し、よく確立されたベンチマークから1500のクエリをキュレートする。
judgeAnythingは、ペア比較とスコア評価の観点から、5段階(GPT-4oやGemini-2.0-Flashなど)の判定能力を評価する。
我々の研究は、より公平な評価プロトコルの必要性と、人間の嗜好との整合性を強調している。
論文 参考訳(メタデータ) (2025-03-21T18:59:20Z) - VisualPRM: An Effective Process Reward Model for Multimodal Reasoning [76.35753243272521]
既存のマルチモーダル大言語モデル(MLLM)の推論能力を改善するVisualPRMを導入する。
我々のモデルは7つのマルチモーダル推論ベンチマークで5.9ポイントの改善を実現している。
マルチモーダルPRMの評価のために,人間に注釈付きステップワイズラベルを付したベンチマークであるVisualProcessBenchを提案する。
論文 参考訳(メタデータ) (2025-03-13T12:03:37Z) - MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models [71.36392373876505]
我々は、LVLM(Large Vision-Language Models)において、インターリーブされたマルチモーダル理解と生成を評価するための大規模ベンチマークであるMMIEを紹介する。
MMIEは、数学、コーディング、物理学、文学、健康、芸術を含む3つのカテゴリ、12のフィールド、102のサブフィールドにまたがる20Kの厳密にキュレートされたマルチモーダルクエリで構成されている。
インターリーブされたインプットとアウトプットの両方をサポートし、多様な能力を評価するために、複数選択とオープンな質問フォーマットの混合を提供する。
論文 参考訳(メタデータ) (2024-10-14T04:15:00Z) - SEED-Bench-2: Benchmarking Multimodal Large Language Models [67.28089415198338]
MLLM(Multimodal large language model)は、最近、テキストだけでなく、インターリーブされたマルチモーダル入力の画像を生成できることを実証した。
SEED-Bench-2は、正確な人間のアノテーションを持つ24Kの多重選択質問で構成されており、27次元にまたがっている。
我々は,23個の著名なオープンソースMLLMの性能を評価し,貴重な観察結果を要約した。
論文 参考訳(メタデータ) (2023-11-28T05:53:55Z) - SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension [27.53415400454066]
生成モデルを評価するためにSEED-Benchというベンチマークを導入する。
SEED-Benchは、正確な人間のアノテーションを持つ19Kの複数の選択質問からなる。
空間的および時間的理解の両面を網羅し,全12次元にわたる18モデルの性能評価を行った。
論文 参考訳(メタデータ) (2023-07-30T04:25:16Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。