論文の概要: P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture
- arxiv url: http://arxiv.org/abs/2606.23256v1
- Date: Mon, 22 Jun 2026 12:38:36 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-24 22:58:51.18005
- Title: P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture
- Title(参考訳): P-JEPA:予測アーキテクチャを組み込んだ手続き型ビデオ表現学習
- Authors: Felix Tristram, Stefano Gasperini, Benjamin Killeen, Marcel Walch, Christian Benz, Nassir Navab, Ghazal Ghazaei,
- Abstract要約: 本稿では,高密度なフレームアラインなアクション空間に問題を還元し,長周期ビデオ表現を学習するバックボーン非依存的手法を提案する。
このアプローチにより、プロシージャ共同埋め込み予測アーキテクチャーは、30分以上のビデオを取り込み、プロシージャステップの効果的なロングフォーム理解を可能にします。
- 参考スコア(独自算出の注目度): 37.47935619712234
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: The increasing maturity of embodied AI platforms has driven a growing interest in procedural video representation learning to support intelligent assistance systems for complex, multi-step tasks. Leveraging large-scale latent predictive training, video foundation models capture video dynamics, enabling downstream tasks such as activity understanding, spatiotemporal localization, and predictive control. However, procedural videos include actions with long-range dependencies that these models do not support, due to the quadratic complexity of self-attention. Distinct actions, for example, may be visually similar despite appearing at different points in the procedure, such as turning the stove on versus off. Here, we propose a backbone-agnostic approach that learns long-duration video representations by reducing the problem to a dense, frame-aligned action space and predicting pooled masked latent vectors. This approach allows our Procedural Joint Embedding Predictive Architecture (P-JEPA) to ingest videos over 30 minutes long, enabling effective long-form understanding of procedural steps. We evaluate P-JEPA using features extracted with VJEPA2.1, TSM, and I3D over the EgoExo4D, EgoProceL, and Assembly101 datasets, finding that it consistently improves linear separability, streaming inference, and temporal action segmentation performance, achieving state-of-the-art results on EgoExo4D fine-grained action classification while using an order of magnitude fewer parameters than LLM-based methods and running in real time.
- Abstract(参考訳): エンボディされたAIプラットフォームの成熟度が高まるにつれ、複雑なマルチステップタスクのためのインテリジェントなアシストシステムをサポートするための手続き型ビデオ表現学習への関心が高まっている。
大規模潜在予測トレーニングを活用することで、ビデオファンデーションモデルは、ビデオダイナミクスをキャプチャし、アクティビティ理解、時空間の局所化、予測制御などの下流タスクを可能にする。
しかし、プロシージャビデオには、これらのモデルがサポートしていない長距離依存のアクションが含まれている。
例えば、特定の動作は、例えばストーブのオン/オフなど、手順の異なる点に現れるにもかかわらず、視覚的に類似している可能性がある。
本稿では,この問題を高密度なフレーム整列アクション空間に還元し,マスク付き潜伏ベクトルのプール化を予測することで,長周期ビデオ表現を学習するバックボーン非依存手法を提案する。
このアプローチにより、プロシージャ共同埋め込み予測アーキテクチャ(P-JEPA)は、30分以上のビデオを取り込み、プロシージャステップの効果的なロングフォーム理解を可能にします。
我々は,VJEPA2.1,TSM,I3DをEgoExo4D,EgoProceL,Ambly101データセット上で抽出した特徴を用いてP-JEPAを評価し,線形分離性,ストリーミング推論,時間的動作セグメンテーション性能を一貫して改善し,EgoExo4Dの粒度の細かい動作分類において,LLM法よりもパラメータの桁数を小さくし,リアルタイムに実行可能であることを発見した。
関連論文リスト
- HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning [0.764671395172401]
Hierarchical Program Probing (HPP) は、長いビデオ理解を階層化されたビデオの反復的、プログラム的な探索として再構成するフレームワークである。
長いビデオ上での探索を可能にするために,情報認識型階層的セグメンテーション,遅延相互作用セマンティック検索,構造化と時間的局所化のための探索機能,という3つのコンポーネントを紹介した。
論文 参考訳(メタデータ) (2026-06-19T20:43:49Z) - InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning [61.87294123934038]
ビデオタスクにおけるマルチモーダル理解を強化するフレームワークであるInternVideo3を提案する。
MCRは理解を、共有、進化するコンテキスト上でのクローズドループプロセスとして扱う。
InternVideo3は、Video-MME、MLVU、Egoなどのベンチマークで強力なパフォーマンスを実現しています。
論文 参考訳(メタデータ) (2026-06-10T15:17:08Z) - Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph [29.737059125885057]
Video-STRは様々なベンチマークで最先端の結果を達成し、ML-Benchではベースモデルを13%上回っている。
コード、モデル、データはリリースされます。
論文 参考訳(メタデータ) (2025-10-13T03:26:56Z) - Spatial-Temporal Multi-level Association for Video Object Segmentation [89.32226483171047]
本稿では,参照フレーム,テストフレーム,オブジェクト特徴を相互に関連付ける空間的・時間的多レベルアソシエーションを提案する。
具体的には,空間的・時間的多段階特徴関連モジュールを構築し,より優れた目標認識特徴を学習する。
論文 参考訳(メタデータ) (2024-04-09T12:44:34Z) - PALM: Predicting Actions through Language Models [74.10147822693791]
本稿では,長期的行動予測の課題に取り組むアプローチであるPALMを紹介する。
本手法は,従来の行動系列を追跡する行動認識モデルと,関連する環境の詳細を記述するための視覚言語モデルを含む。
実験の結果,PALMは長期的な行動予測作業において最先端の手法を超越していることがわかった。
論文 参考訳(メタデータ) (2023-11-29T02:17:27Z) - Temporal DINO: A Self-supervised Video Strategy to Enhance Action
Prediction [15.696593695918844]
本稿では、DINOにインスパイアされた行動予測(ラベルのない自己蒸留)を強化するための、新しい自己教師型ビデオ戦略を提案する。
実験結果は、3D-ResNet、Transformer、LSTMアーキテクチャで予測性能が大幅に向上したことを示している。
これらの知見は,行動認識,運動計画,シーン理解など,多様な映像ベースタスクにおけるアプローチの可能性を強調した。
論文 参考訳(メタデータ) (2023-08-08T21:18:23Z) - Leaping Into Memories: Space-Time Deep Feature Synthesis [93.10032043225362]
内部モデルから映像を合成するアーキテクチャ非依存の手法であるLEAPSを提案する。
我々は,Kineetics-400に基づく多種多様なアーキテクチャの進化的注目を反転させることにより,LEAPSの適用性を定量的かつ定性的に評価する。
論文 参考訳(メタデータ) (2023-03-17T12:55:22Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。