論文の概要: Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
- arxiv url: http://arxiv.org/abs/2604.05015v1
- Date: Mon, 06 Apr 2026 17:59:56 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-04-08 17:42:09.409859
- Title: Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
- Title(参考訳): Video-MME-v2: 総合的ビデオ理解のためのベンチマークの次の段階に向けて
- Abstract要約: Video-MME-v2は、ビデオ理解の堅牢性と忠実さを厳格に評価するために設計された総合的なベンチマークである。
データ品質を保証するため、Video-MME-v2は厳格に制御された人間のアノテーションパイプラインを通して構築される。
- 参考スコア(独自算出の注目度): 98.3098451637867
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: With the rapid advancement of video understanding, existing benchmarks are becoming increasingly saturated, exposing a critical discrepancy between inflated leaderboard scores and real-world model capabilities. To address this widening gap, we introduce Video-MME-v2, a comprehensive benchmark designed to rigorously evaluate the robustness and faithfulness of video understanding. To systematically evaluate model capabilities, we design a \textbf{progressive tri-level hierarchy} that incrementally increases the complexity of video comprehension, ranging from multi-point visual information aggregation, to temporal dynamics modeling, and ultimately to complex multimodal reasoning. Besides, in contrast to conventional per-question accuracy, we propose a \textbf{group-based non-linear evaluation} strategy that enforces both consistency across related queries and coherence in multi-step reasoning. It penalizes fragmented or guess-based correctness and assigns credit only to answers supported by valid reasoning. To guarantee data quality, Video-MME-v2 is constructed through a rigorously controlled human annotation pipeline, involving 12 annotators and 50 independent reviewers. Backed by \textbf{3,300 human-hours} and up to \textbf{5 rounds} of quality assurance, Video-MME-v2 aims to serve as one of the most authoritative video benchmarks. Extensive experiments reveal a substantial gap between current best model Gemini-3-Pro and human experts, and uncover a clear hierarchical bottleneck where errors in visual information aggregation and temporal modeling propagate to limit high-level reasoning. We further find that thinking-based reasoning is highly dependent on textual cues, improving performance with subtitles but sometimes degrading it in purely visual settings. By exposing these limitations, Video-MME-v2 establishes a demanding new testbed for the development of next-generation video MLLMs.
- Abstract(参考訳): ビデオ理解の急速な進歩により、既存のベンチマークは飽和し、膨らませたリーダーボードスコアと現実世界のモデル能力の間に重要な違いが浮かび上がっている。
この拡張ギャップに対処するため,ビデオ理解の堅牢性と忠実さを徹底的に評価するための総合的なベンチマークであるVideo-MME-v2を導入する。
モデル機能を体系的に評価するために,多点視覚情報集約から時間的ダイナミクスモデリング,そして最終的には複雑なマルチモーダル推論に至るまで,ビデオ理解の複雑さを漸進的に増大させる「textbf{progressive tri-level hierarchy」を設計する。
また,従来の問合せ精度とは対照的に,関連するクエリ間の一貫性と複数ステップの推論におけるコヒーレンスを両立させる,textbf{group-based non-linear evaluation} 戦略を提案する。
断片的または推測に基づく正しさを罰し、有効な推論によってサポートされている回答にのみクレジットを割り当てる。
データ品質を保証するため、Video-MME-v2は、12のアノテーションと50の独立したレビュアーを含む厳格に制御された人間のアノテーションパイプラインを通して構築される。
品質保証の \textbf{3,300 human-hours} と \textbf{5 rounds} に支援された Video-MME-v2 は、最も権威あるビデオベンチマークの1つとして機能することを目指している。
大規模な実験では、現在の最良のモデルであるGemini-3-Proと人間の専門家の間に大きなギャップが見られ、視覚情報集約と時間的モデリングにおけるエラーが、高レベルの推論を制限するために伝播する明確な階層的ボトルネックが明らかになった。
さらに、思考に基づく推論はテキストの手がかりに大きく依存しており、字幕によるパフォーマンスを向上させるが、純粋に視覚的な設定では劣化することがある。
これらの制限を明らかにすることで、ビデオMME-v2は次世代のビデオMLLMの開発に必要な新しいテストベッドを確立する。
関連論文リスト
- Kwai Keye-VL-2.0 Technical Report [53.82434681649277]
Keye-VL-2.0は、長期ビデオ理解とエージェントインテリジェンスを促進するために設計されたマルチモーダル基盤モデルである。
DeepSeek Sparse Attention (DSA)をGQAベースのマルチモーダルアーキテクチャに適応したのは,これが初めてである。
コンテクスト-RLとビデオ-RLを併用したMOPD(Cross-Modal Multi-Teacher On-Policy Distillation)は破滅的忘れのアルゴリズム的ジレンマを克服する。
論文 参考訳(メタデータ) (2026-06-09T09:58:08Z) - HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding [1.8482679687103294]
HAVENは階層的に整列したマルチモーダル・ベンチマークである。
この統一アノテーションパラダイムに基づいて,要約,時間的推論,マルチモーダルグラウンド,サリエンシランキングにまたがる総合評価スイートを提案する。
論文 参考訳(メタデータ) (2026-05-19T00:48:14Z) - MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation [48.84450712826316]
MSVBenchは、マルチショットビデオ生成に適した階層的なスクリプトと参照イメージを備えた最初の包括的なベンチマークである。
本稿では,大規模マルチモーダルモデルの高レベルな意味推論と,ドメイン固有のエキスパートモデルの微粒な知覚的厳密さを相乗化するハイブリッド評価フレームワークを提案する。
論文 参考訳(メタデータ) (2026-02-27T12:26:34Z) - UI2V-Bench: An Understanding-based Image-to-video Generation Benchmark [35.157850129371525]
I2V(Image-to-Video)の生成は、ビデオ合成の分野において重要な焦点となっている。
既存の評価ベンチマークは主にビデオの品質や時間的一貫性といった側面に焦点を当てている。
We propose UI2V-Bench, a novel benchmark for evaluation I2V model with focus on semantic understanding and reasoning。
論文 参考訳(メタデータ) (2025-09-29T08:14:26Z) - VideoScore2: Think before You Score in Generative Video Evaluation [69.43069741467603]
VideoScore2は、視覚的品質、テキスト・ツー・ビデオのアライメント、物理的/常識的一貫性を明確に評価する多次元、解釈可能、そして人間によるアライメントフレームワークである。
我々のモデルは、27,168人の注釈付きビデオを含む大規模なデータセットVideoFeedback2で訓練されている。
論文 参考訳(メタデータ) (2025-09-26T18:09:03Z) - HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding [120.84817886550765]
MLLM(Multimodal Large Language Models)は、画像とビデオの両方を含む視覚的理解タスクにおいて、大きな進歩を見せている。
既存の人間中心のベンチマークは、主にビデオ生成の品質と行動認識を強調し、人間中心のシナリオに必要な知覚と認知の能力を見落としている。
我々は,人間中心のビデオ理解におけるMLLMのより総合的な評価を提供するために,厳格にキュレートされたベンチマークを提案する。
論文 参考訳(メタデータ) (2025-07-07T11:52:24Z) - Bridging Video Quality Scoring and Justification via Large Multimodal Models [14.166920184033463]
古典的映像品質評価法(VQA)は、映像の視覚的忠実さと明瞭さを判断する数値スコアを生成する。
しかし、スコアはビデオの複雑な品質の次元を表現できず、適用性を制限する。
言語出力から恩恵を受け、ビデオ大マルチモーダルモデル(LMM)を命令チューニングによりVQAに適応させることは、この問題に対処する可能性がある。
論文 参考訳(メタデータ) (2025-06-26T05:02:25Z) - VideoMolmo: Spatio-Temporal Grounding Meets Pointing [66.19964563104385]
VideoMolmoは、ビデオシーケンスのきめ細かいポインティングに適したモデルだ。
新しい仮面融合はSAM2を双方向の点伝播に用いている。
The generalization of VideoMolmo, we introduced VPoMolS-temporal, a challenge out-of-distribution benchmark across two real-world scenarios。
論文 参考訳(メタデータ) (2025-06-05T17:59:29Z) - H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding [25.111988967973147]
既存のビデオ理解評価ベンチマークでは、カバレッジ、タスクの多様性、シーン適応性に大きな制限がある。
本稿では,一般的なビデオとオンラインストリーミングの両方の理解度を評価するために,階層的・全体論的ビデオ理解ベンチマークを提案する。
このベンチマークは、拡張ビデオの長さ、包括的なアセスメントタスク、エンリッチ化ビデオデータという3つの重要な特徴に寄与する。
論文 参考訳(メタデータ) (2025-03-31T12:32:51Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。