論文の概要: K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos
- arxiv url: http://arxiv.org/abs/2607.02680v1
- Date: Thu, 02 Jul 2026 18:23:14 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-07 22:26:29.39014
- Title: K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos
- Title(参考訳): K9-Bench: 犬中心ビデオにおけるマルチモーダルLCMの評価
- Authors: Khush Attarde, Yusuf Ali, Megha Thukral, Divye Bhutani, Thomas Ploetz, Zsolt Kira,
- Abstract要約: K9-Benchは、実世界のドッグビデオに焦点を当てた、新しいベンチマークである。
犬中心のビデオを自動的にWebからマイニングする,スケーラブルなVLM/LLMベースのデータ生成パイプラインを提案する。
我々は、データセットキュレーション中にVLMが導入したバイアスを排除するために、バイアス軽減戦略を実装した。
- 参考スコア(独自算出の注目度): 24.300198465747044
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: MLLMs have shown strong zero-shot capabilities across diverse inputs such as across images, video, audio, and text. A crucial, yet underexplored, application of these models lies in understanding and modeling animal-centric scenarios. As animals are integral to millions of households, benchmarking next-generation AI models on pet-focused tasks, ranging from recognizing distress signals to enabling responsive robotic companions, is essential for building AI systems that can work alongside us. We introduce K9-Bench, a novel benchmark focused on real-world domestic dog videos, specifically targeting canine action and interaction understanding via approximately 5000 question-answer pairs across 907 videos spanning 5 distinct task categories that test long-form, canine-centric multimodal reasoning in MLLMs. To create this dataset, we propose a scalable, VLM/LLM-powered data generation pipeline that automatically mines canine-centric videos from the web and curates QA pairs requiring fine-grained, multi-hop reasoning over canine actions and temporally extended interaction sequences. We implement bias mitigation strategies designed to eliminate biases introduced by VLMs during dataset curation. Through extensive experimentation, we find that frontier MLLMs exhibit limited zero-shot performance on canine-centric tasks: although state-of-the-art closed-source models outperform open-source counterparts, they still struggle with compositional reasoning over subtle posture and interaction cues spread over long horizons. We observe that generic chain-of-thought prompting provides only modest performance for such long-horizon reasoning. Beyond a novel dataset for canine activity analysis, K9-Bench provides a general-purpose dataset construction pipeline that can be adapted to other low-data domains for quantitative analysis. Our project website is available at: https://ogmenrobotics.github.io/K9Bench.
- Abstract(参考訳): MLLMは画像、ビデオ、オーディオ、テキストなど多様な入力にまたがる強力なゼロショット機能を示している。
これらのモデルの重要かつ過小評価されている応用は、動物中心のシナリオを理解し、モデル化することにある。
動物は何百万もの家庭に不可欠なので、ペットに焦点を当てたタスクで次世代AIモデルをベンチマークする。
我々はK9-Benchを紹介した。K9-Benchは、実世界の犬ビデオに焦点を当てた新しいベンチマークで、犬行動と相互作用の理解を、MLLMの長めの犬中心のマルチモーダル推論をテストする5つのタスクカテゴリにまたがる約5000の質問応答ペアを用いて、特にターゲットとしている。
このデータセットを作成するために、VLM/LLMを利用したスケーラブルなデータ生成パイプラインを提案し、Webから犬中心のビデオを自動的にマイニングし、犬行動と時間的に拡張されたインタラクションシーケンスに対して、細粒度でマルチホップな推論を必要とするQAペアをキュレートする。
我々は、データセットキュレーション中にVLMが導入したバイアスを排除するために、バイアス軽減戦略を実装した。
最先端のクローズドソースモデルはオープンソースモデルよりも優れているが、それでも微妙な姿勢や相互作用の手がかりが長い地平線を越えて広がるという構成的推論に苦慮している。
我々は、このような長い水平推論に対して、ジェネリック・チェーン・オブ・シークレット・プロンプトは控えめな性能しか提供しないのを観察する。
犬の活動分析のための新しいデータセットの他に、K9-Benchは、定量分析のために他の低データドメインに適応可能な汎用データセット構築パイプラインを提供する。
プロジェクトのWebサイトは、https://ogmenrobotics.github.io/K9Bench.comで公開されている。
関連論文リスト
- InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning [61.87294123934038]
ビデオタスクにおけるマルチモーダル理解を強化するフレームワークであるInternVideo3を提案する。
MCRは理解を、共有、進化するコンテキスト上でのクローズドループプロセスとして扱う。
InternVideo3は、Video-MME、MLVU、Egoなどのベンチマークで強力なパフォーマンスを実現しています。
論文 参考訳(メタデータ) (2026-06-10T15:17:08Z) - CVBench: Evaluating Cross-Video Synergies for Complex Multimodal Understanding and Reasoning [11.478276629279526]
CVBenchは,ビデオ間のリレーショナル推論を厳格に評価するために設計された,最初の総合的なベンチマークである。
CVBenchは、クロスビデオオブジェクトアソシエーション、クロスビデオイベントアソシエーション、クロスビデオ複合推論の3層にまたがる1000の質問応答ペアで構成されている。
5つのドメインの異なるビデオクラスタから構築されたこのベンチマークは、ダイナミックな視覚的コンテキストにまたがる情報を合成するモデルに挑戦する。
論文 参考訳(メタデータ) (2025-08-27T03:29:35Z) - EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering [59.94048858464922]
EgoCrossは、EgocentricQAにおけるMLLMのクロスドメイン一般化を評価するためのベンチマークである。
EgoCrossは、手術、産業、極端なスポーツ、動物の観点からの4つの分野をカバーしている。
798のビデオクリップにまたがる約1000のQAペアで構成され、予測、認識、ローカライゼーション、カウントという4つの重要なQAタスクにまたがる。
論文 参考訳(メタデータ) (2025-08-14T15:11:20Z) - Act-as-Pet: Benchmarking the Abilities of Large Language Models as E-Pets in Social Network Services [23.84256497132106]
本稿では,Large Language Models (LLM) を評価するベンチマークであるPet-Benchを紹介する。
Pet-Bench氏は、対話的なエンゲージメントとともに、自己進化と発達の振る舞いを強調し、ペットの仲間関係をよりリアルに反映している。
論文 参考訳(メタデータ) (2025-06-04T09:25:52Z) - CinePile: A Long Video Question Answering Dataset and Benchmark [55.30860239555001]
我々は、CinePileという新しいデータセットとベンチマークを提示する。
包括的データセットは305,000の多重選択質問(MCQ)から構成されており、様々な視覚的・マルチモーダル的な側面をカバーしている。
トレーニングスプリットに関して、オープンソースのVideo-LLMを微調整し、データセットのテストスプリット上で、オープンソースとプロプライエタリなビデオ中心LLMの両方を評価しました。
論文 参考訳(メタデータ) (2024-05-14T17:59:02Z) - How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs [98.37571997794072]
CVRR-ES(Complex Video Reasoning and Robustness Evaluation Suite)について紹介する。
CVRR-ESは、11種類の実世界のビデオ次元にわたるビデオLMMの性能を包括的に評価する。
我々の発見は、次世代の人間中心AIシステムを構築する上で貴重な洞察を提供する。
論文 参考訳(メタデータ) (2024-05-06T17:59:45Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。