論文の概要: RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail
- arxiv url: http://arxiv.org/abs/2607.00310v1
- Date: Wed, 01 Jul 2026 01:23:34 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-02 19:56:07.677943
- Title: RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail
- Title(参考訳): RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail
- Authors: Amirreza Rouhi, Rajat Aggarwal, Parikshit Sakurikar, Anoop M. Namboodiri, Sashi P. Reddi,
- Abstract要約: 本研究では,事前学習された基礎的ビデオワールドモデルの小売シーンへのパラメータ効率の適応について検討する。
店舗・店舗の視点から, 5つのスーパーマーケットから32,105個のキャプション付き小売クリップのコーパスであるRetailSMVを紹介した。
また, LPIPS, PSNR, ドリームシムでは, エキソセンタのみの適応が一致し, 組み合わせた適応が7点のうち6点に留まり, LPIPS, PSNR, ドリームシムの方が有意に優れていた。
- 参考スコア(独自算出の注目度): 1.8594711725515678
- License: http://creativecommons.org/licenses/by-nc-sa/4.0/
- Abstract: Foundation video diffusion models are increasingly viewed as world simulators for embodied agents, yet their pretraining on internet-scale generic video leaves them poorly aligned with real-world deployment domains. We study parameter-efficient adaptation of a pretrained foundation video world model to retail scenes: when synchronized egocentric and exocentric video of the same activity are available, which viewpoint of training data produces the strongest adapted model? We introduce RetailSMV (Retail Synchronized Multi-View), a corpus of 32,105 captioned retail clips from five supermarkets with synchronized ego/exo capture from the store-staff perspective (stocking, arranging, weighing, managing supply carts, scanning at checkout), rather than the customer-centric framing of prior retail video corpora, and train three matched Low-Rank Adaptation (LoRA) configurations of Cosmos3-Nano (egocentric-only, exocentric-only, combined) under identical hyperparameters. On a 200-clip held-out test set evaluated with seven complementary metrics under a strict paired statistical protocol, exocentric-only adaptation matches or exceeds combined adaptation on six of seven point estimates and is significantly better on LPIPS, PSNR, and DreamSim, despite training on only 15,985 exocentric clips (versus 32,105 for combined). A symmetric paired comparison further shows that adding exocentric data to egocentric-only training helps while adding egocentric data to exocentric-only training hurts. The absolute adaptation gap is largest at the shortest rollout time, identifying the near-horizon prediction window as the regime in which adaptation is most beneficial.
- Abstract(参考訳): 基礎的なビデオ拡散モデルは、エンボディエージェントのワールドシミュレータとしてますます見なされているが、インターネットスケールのジェネリックビデオでの事前訓練は、現実のデプロイメントドメインとの整合性を損なう。
我々は、事前学習された基礎的ビデオワールドモデルの小売分野への適応性について検討する:同一活動の共生型エゴセントリック・エキソセントリックビデオが利用可能である場合、トレーニングデータのどの視点が最強適応モデルを生成するか。
RetailSMV (Retail Synchronized Multi-View) は、小売ビデオコーパスの顧客中心のフレーミングではなく、店員視点(ストック、アレンジ、アレンジ、サプライカートの管理、チェックアウト)から抽出した5つのスーパーマーケットからの32,105個のキャプション付き小売クリップのコーパスと、コスモス3-ナノ(エゴ中心、エゴ中心、エクソ中心、複合)の3つのマッチしたローランク適応(LoRA)構成を同一のハイパーパラメータで導入する。
厳密なペア付き統計プロトコルの下で7つの相補的な指標で評価された200クリップのホールトアウトテストセットでは、外心のみの適応一致が7つの評価値のうち6つの評価値で組み合わせられ、LPIPS、PSNR、DreamSimでかなり優れている(合計32,105回)。
対称対比較により、エゴセントリックなトレーニングにエゴセントリックなデータを追加することは、エゴセントリックなトレーニングにエゴセントリックなデータを加えるのに役立ちます。
絶対適応ギャップは、最も短いロールアウト時間で最大であり、ほぼ水平の予測ウィンドウを、適応が最も有益であるレジームとして識別する。
関連論文リスト
- EgoExo-WM: Unlocking Exo Video for Ego World Models [59.337519140691775]
エゴセントリックな世界モデルはエージェントの予測と計画を可能にする有望な方向性を示す。
外心ビデオは豊富で、身体のポーズは良好だが、エージェントのアクション空間と直接的に一致していない。
本稿では,このギャップを埋めるために,アクションの表現として,外心ビデオから構造体ポーズを抽出する手法を提案する。
論文 参考訳(メタデータ) (2026-05-14T23:35:54Z) - WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation [51.1909041777449]
We present WorldWander, a in-context learning framework designed for translating between egocentric and exocentric worlds in video generation。
実験により、WorldWanderは優れた視点同期、文字一貫性、一般化を実現している。
論文 参考訳(メタデータ) (2025-11-27T04:40:37Z) - Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning [80.37314291927889]
EMBEDは、エゴセントリックなビデオ表現学習のための、エゴセントリックなビデオ言語データを変換するために設計された手法である。
エゴセントリックなビデオは、主にクローズアップなハンドオブジェクトのインタラクションを特徴としているのに対し、エゴセントリックなビデオは、人間の活動に対してより広い視点を提供する。
視覚と言語スタイルの転送の両方を適用することで、私たちのフレームワークは新しいエゴセントリックなデータセットを作成します。
論文 参考訳(メタデータ) (2024-08-07T06:10:45Z) - Ego-VPA: Egocentric Video Understanding with Parameter-efficient Adaptation [57.38965505987893]
Ego-VPAは、エゴ中心のビデオタスクに対するパラメータ効率の適応である。
Ego-VPAは、わずか0.84%の学習可能なパラメータで軽量な適応を実現している。
論文 参考訳(メタデータ) (2024-07-28T16:01:32Z) - EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding [27.881857222850083]
EgoExo-Fitnessは新しいフルボディアクション理解データセットである。
シンクロナイズドエゴセントリックカメラと固定型エゴセントリックカメラで撮影されたフィットネス・シーケンス・ビデオが特徴。
EgoExo-Fitnessは、エゴセントリックでエゴセントリックなフルボディの行動理解を研究するための新しいリソースを提供する。
論文 参考訳(メタデータ) (2024-06-13T07:28:45Z) - X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization [56.75782714530429]
我々はX-MICと呼ぶクロスモーダル適応フレームワークを提案する。
私たちのパイプラインは、凍結したテキストの埋め込みを、共有された埋め込み空間内で、それぞれのエゴセントリックなビデオにアライメントすることを学びました。
これにより、各エゴセントリックビデオへのテキスト埋め込みのアライメントが向上し、データセットの一般化が大幅に向上する。
論文 参考訳(メタデータ) (2024-03-28T19:45:35Z) - Synchronization is All You Need: Exocentric-to-Egocentric Transfer for Temporal Action Segmentation with Unlabeled Synchronized Video Pairs [14.61648563523105]
我々は、当初、外向型(固定型)カメラ用に設計された時間的アクションセグメンテーションシステムを、ウェアラブルカメラが映像データをキャプチャするエゴセントリックなシナリオに転送する問題を考える。
本稿では,既存のラベル付きエキソセントリックビデオを活用する新しい手法と,ラベル付き,同期化されたエキソセントリックビデオペアの新たなセットを提案する。
Assembly101とEgoExo4Dの実験は、従来の教師なし領域適応と時間的アライメントに対する提案手法の有効性を示した。
論文 参考訳(メタデータ) (2023-12-05T10:24:43Z) - Ego-Exo: Transferring Visual Representations from Third-person to
First-person Videos [92.38049744463149]
大規模第3者映像データセットを用いた自己中心型映像モデルの事前訓練手法について紹介する。
私たちのアイデアは、重要なエゴセントリック特性を予測する第三者ビデオから潜在信号を見つけることです。
実験の結果,Ego-Exoフレームワークは標準ビデオモデルにシームレスに統合可能であることがわかった。
論文 参考訳(メタデータ) (2021-04-16T06:10:10Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。