論文の概要: Watch Before You Answer: Learning from Visually Grounded Post-Training
- arxiv url: http://arxiv.org/abs/2604.05117v1
- Date: Mon, 06 Apr 2026 19:22:48 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-04-08 17:42:09.463694
- Title: Watch Before You Answer: Learning from Visually Grounded Post-Training
- Title(参考訳): 振り返る前に見る: 視界から学ぶ
- Authors: Yuxuan Zhang, EunJeong Hwang, Huaisong Zhang, Penghui Du, Yiming Jia, Dongfu Jiang, Xuan He, Shenhui Zhang, Ping Nie, Peter West, Kelsey R. Allen,
- Abstract要約: ビデオ理解のパフォーマンスは、まだテキストベースの推論に遅れている。
一般的に報告されているベンチマークには、テキストキューだけで答えられる40~60%の質問が含まれている。
VidGroundは、視覚的に接地された質問のみを用いて、シンプルで効果的なソリューションとして紹介する。
- 参考スコア(独自算出の注目度): 27.736034325218554
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: It is critical for vision-language models (VLMs) to comprehensively understand visual, temporal, and textual cues. However, despite rapid progress in multimodal modeling, video understanding performance still lags behind text-based reasoning. In this work, we find that progress is even worse than previously assumed: commonly reported long video understanding benchmarks contain 40-60% of questions that can be answered using text cues alone. Furthermore, we find that these issues are also pervasive in widely used post-training datasets, potentially undercutting the ability of post-training to improve VLM video understanding performance. Guided by this observation, we introduce VidGround as a simple yet effective solution: using only the actual visually grounded questions without any linguistic biases for post-training. When used in tandem with RL-based post-training algorithms, this simple technique improves performance by up to 6.2 points relative to using the full dataset, while using only 69.1% of the original post-training data. Moreover, we show that data curation with a simple post-training algorithm outperforms several more complex post-training techniques, highlighting that data quality is a major bottleneck for improving video understanding in VLMs. These results underscore the importance of curating post-training data and evaluation benchmarks that truly require visual grounding to advance the development of more capable VLMs. Project page: http://vidground.etuagi.com.
- Abstract(参考訳): 視覚言語モデル(VLM)は、視覚的、時間的、テキスト的な手がかりを包括的に理解することが重要である。
しかし、マルチモーダルモデリングの急速な進歩にもかかわらず、ビデオ理解性能はテキストベースの推論よりも遅れている。
一般的に報告されている長いビデオ理解ベンチマークには、テキストキューだけで答えられる40~60%の質問が含まれている。
さらに、これらの問題はトレーニング後データセットでも広く利用されており、VLMビデオ理解性能を向上させるためのトレーニング後の能力を弱めている可能性がある。
この観察に導かれ、我々はVidGroundをシンプルで効果的なソリューションとして紹介した。
RLベースのポストトレーニングアルゴリズムと組み合わせて使用すると、この単純なテクニックは、元のトレーニング後のデータの69.1%しか使用せず、完全なデータセットの使用に対して最大6.2ポイントの性能を向上させる。
さらに、単純な後学習アルゴリズムによるデータキュレーションは、VLMにおけるビデオ理解を改善する上で、データ品質が大きなボトルネックであることを強調し、より複雑な後学習手法よりも優れていることを示す。
これらの結果は、より有能なVLMの開発を進めるために、ビジュアルグラウンドティングを本当に必要とするような、トレーニング後のデータと評価ベンチマークのキュレーションの重要性を浮き彫りにしている。
プロジェクトページ: http://vidground.etuagi.com
関連論文リスト
- TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs [81.78017865436816]
我々は,映像の時間的接地能力の強いMLLMを体系的に構築するTimeLensを提案する。
まず,既存のVTGベンチマークにおける重要な品質問題を明らかにし,TimeLens-Benchを導入する。
また、自動再アノテーションパイプラインを通じてノイズの多いトレーニングデータに対処し、大規模で高品質なトレーニングデータセットであるTimeLens-100Kを出力します。
論文 参考訳(メタデータ) (2025-12-16T18:59:58Z) - Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding [57.26400319795876]
時間的ビデオグラウンディング(TVG)は、長めのビデオ理解における中核的な課題である。
近年のLVLM(Large Vision-Language Models)は,教師付き微調整によるTVG処理の早期実現を示唆している。
強化学習によるLVLMの一般化能力を高める新しいポストトレーニングフレームワークを提案する。
論文 参考訳(メタデータ) (2025-03-17T17:04:20Z) - FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability [10.184567639685321]
本稿では,LVLMを学習するための新しいデータセット構築手法であるFiVLを紹介する。
本稿では,モデルがイメージを実体的証拠として用いる能力を評価するためのベンチマークを示す。
視覚による幻覚を説明できる最強の視覚言語アライメントで注目頭を特定する。
論文 参考訳(メタデータ) (2024-12-19T09:24:10Z) - Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation [57.34255010956452]
この研究は、合成データによるスケーリングを再考し、データ中心の観点からビデオLLMの開発に焦点を当てる。
本研究では,純粋なテキスト命令データからビデオライクなサンプルを合成するSparrowというデータ拡張手法を提案する。
提案手法は,より多くのサンプルを用いてトレーニングしたベースラインに匹敵する,あるいは優れた性能を実現する。
論文 参考訳(メタデータ) (2024-11-29T18:59:54Z) - PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning [78.23573511641548]
視覚言語事前学習は、幅広い画像言語アプリケーションで性能を大幅に向上させた。
しかし、ビデオ関連タスクの事前学習プロセスは、非常に大きな計算とデータリソースを必要とする。
本稿では,映像理解のための既存の画像言語事前学習モデルに適用するための,ストレートフォワード,高効率,資源光のアプローチについて検討する。
論文 参考訳(メタデータ) (2024-04-25T19:29:55Z) - CUPID: Adaptive Curation of Pre-training Data for Video-and-Language
Representation Learning [49.18591896085498]
ソースデータとターゲットデータのドメインギャップを埋めるCUPIDを提案します。
CUPIDは、複数のビデオ言語およびビデオタスクにまたがる最新のパフォーマンスを提供します。
論文 参考訳(メタデータ) (2021-04-01T06:42:16Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。