論文の概要: Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos
- arxiv url: http://arxiv.org/abs/2607.11523v1
- Date: Mon, 13 Jul 2026 13:09:45 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-14 17:47:21.479058
- Title: Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos
- Title(参考訳): Vinci2: 連続エゴセントリックビデオにおける積極的なアシストを提供する
- Authors: Gong Sitong, Tianyu Yan, Caixin Kang, Bo Zheng, Xiang Ruan, Huchuan Lu, Kaipeng Zhang, Yoichi Sato, Yifei Huang,
- Abstract要約: 継続的エゴセントリックなビデオは、リッチで進化したコンテキストを提供し、新しい形式の支援を可能にします。
Vinci2は、デバイス上のアシスタントであるVinciを、反応反応から活性へ前進させる、活性中心型補助システムである。
EgoMemoはトレーニング不要で、メモリ拡張されたエージェントで、3つの補完的なメモリ表現を保持する。
- 参考スコア(独自算出の注目度): 74.71281081304846
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: When should an intelligent assistant speak up without being asked? Continuous egocentric video offers rich, evolving context that enables a new form of assistance: one that is proactive rather than merely reactive. Yet existing approaches either wait passively for user queries or treat every detected event as requiring a response, without considering the user's history, current activity, or whether assistance would actually be welcome. We reframe proactive assistance as a context-dependent decision problem: the agent must not only perceive what is happening, but reason over accumulated temporal context to determine when and whether to intervene. To this end, we present Vinci2, a proactive egocentric assistance system that advances the on-device assistant Vinci from reactive response toward proactivity. On the evaluation side, we present EgoServe, the first large-scale benchmark for proactive assistance in continuous egocentric video. EgoServe comprises over 3,000 service instances organized along 4 temporal memory horizons, ranging from immediate safety alerts to long-term habit coaching, across 10 service categories. On the modeling side, we propose EgoMemo, a training-free, memory-augmented agent that maintains three complementary memory representations: multi-scale temporal summaries, a semantic knowledge graph, and visual embedding archives. At each timestep, EgoMemo performs retrieval-augmented reasoning to determine whether assistance is warranted and, if so, produces contextually grounded responses. Experiments demonstrate that EgoMemo establishes strong baselines on EgoServe while remaining competitive on existing egocentric benchmarks. Our benchmark and code are publicly available at \href{https://sitonggong.github.io/EgoServe-page/}{Vinci2}.
- Abstract(参考訳): インテリジェントアシスタントはいつ、尋ねられることなく話すべきか?
継続的エゴセントリックなビデオは、リッチで進化したコンテキストを提供し、新しい形式の支援を可能にします。
しかし、既存のアプローチでは、ユーザクエリを受動的に待機するか、検出されたすべてのイベントを、ユーザの履歴や現在のアクティビティ、あるいは実際にアシストが歓迎されるかどうかを考慮せずに、応答を必要とするものとして扱う。
エージェントは、何が起こっているかだけでなく、蓄積した時間的コンテキストを過度に理解して、いつ、いつ、介入すべきかを判断しなければなりません。
この目的のために、デバイス上のアシスタントであるVinciを反応応答から活性へ前進させる、活性中心型補助システムであるVinci2を提案する。
評価面では、エゴセントリックビデオにおけるプロアクティブアシストのための最初の大規模ベンチマークであるEgoServeを紹介する。
EgoServeは4つの時間的メモリ水平線に沿って構成された3000以上のサービスインスタンスで構成されている。
モデリング面では,マルチスケールの時間的要約,意味知識グラフ,視覚的埋め込みアーカイブという,3つの相補的なメモリ表現を保持する,トレーニングフリーでメモリ拡張されたエージェントであるEgoMemoを提案する。
それぞれの段階において、EgoMemoは、補助が保証されているかどうかを判断するために、検索強化された推論を実行し、もしそうであれば、文脈的に根拠づけられた応答を生成する。
実験によると、EgoMemoはEgoServeの強力なベースラインを確立しつつ、既存のエゴセントリックなベンチマークで競争力を維持している。
私たちのベンチマークとコードは、 \href{https://sitonggong.github.io/EgoServe-page/}{Vinci2}で公開されています。
関連論文リスト
- EgoIntrospect: An Egocentric Dataset and Benchmark for User-Centric Internal State Reasoning [47.853306116245484]
EgoIntrospectは、セルフアノテーションを備えたユーザ駆動のシナリオでキャプチャされた最初のエゴセントリックなデータセットである。
収録時間は60人から180時間、平均録音時間は1人あたり3時間である。
我々は、感情経験、インタラクティブな意図、認知記憶など、ユーザ内部状態を中心とした一連のタスクを形式化する。
論文 参考訳(メタデータ) (2026-05-17T05:05:29Z) - EgoLife: Towards Egocentric Life Assistant [60.51196061794498]
我々はEgoLifeを紹介した。EgoLifeは、AIを使ったウェアラブルグラスを通じて、個人の効率を向上するエゴセントリックなライフアシスタントを開発するプロジェクトだ。
我々は、6人の参加者が1週間一緒に暮らし、マルチモーダル・エゴセントリックなビデオキャプチャーにAIグラスを使用して日々の活動を継続的に記録し、同期された3人称ビデオ参照を行う総合的なデータ収集研究を行った。
この取り組みの結果、EgoLifeデータセットは、集中的なアノテーションを備えた300時間のエゴセントリック、対人、マルチビュー、マルチモーダルの日常生活データセットである。
私たちはEgoLifeQAを紹介します。EgoLifeQAは、長いコンテキスト、ライフ指向の質問応答タスクのスイートで、提供するように設計されています。
論文 参考訳(メタデータ) (2025-03-05T18:54:16Z) - EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering [95.2396264550978]
シーンテキストを含むエゴセントリックなQA支援のための,斬新で厳密に構築されたベンチマークであるEgoTextVQAを紹介する。
EgoTextVQAには1.5Kのエゴビュービデオと7Kのシーンテキスト対応の質問が含まれており、屋外運転や屋内ホームキーピング活動における実際のユーザニーズを反映している。
論文 参考訳(メタデータ) (2025-02-11T09:45:06Z) - Egocentric and Exocentric Methods: A Short Survey [25.41820386246096]
エゴセントリックな視覚は、カメラ装着者の視点からシーンを捉えます。
外見中心の視覚はシーン全体のコンテキストを捉えます。
エゴとエクソビューの併用モデリングは、次世代AIエージェントの開発に不可欠である。
論文 参考訳(メタデータ) (2024-10-27T22:38:51Z) - Retrieval-Augmented Egocentric Video Captioning [53.2951243928289]
EgoInstructor(エゴインストラクタ)は、意味的に関連する第三者の指導ビデオを自動的に検索する、検索拡張マルチモーダルキャプションモデルである。
我々は、エゴセントリックでエゴセントリックなビデオ機能を引き出す新しいEgoExoNCE損失でクロスビュー検索モジュールをトレーニングし、同様のアクションを記述した共有テキスト機能にアライメントすることで、より近づいた。
論文 参考訳(メタデータ) (2024-01-01T15:31:06Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。