論文の概要: How You Move Tells What You'll Do: Trajectory-Conditioned Egocentric Prediction
- arxiv url: http://arxiv.org/abs/2605.20388v1
- Date: Tue, 19 May 2026 18:38:11 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-05-21 19:19:56.324735
- Title: How You Move Tells What You'll Do: Trajectory-Conditioned Egocentric Prediction
- Title(参考訳): 軌跡に基づくエゴセントリック予測(動画あり)
- Authors: Sejoon Jun, Hai Nguyen-Truong, Luigi Seminara, Lorenzo Torresani,
- Abstract要約: 我々はこれらの知見を,エゴセントリックな文脈から将来の軌道候補を予測するモデルであるTrajPilotとしてインスタンス化する。
TrajPilotは、Ego-Exo4D Atomic、Ego-Exo4D Keystep、Ego4D GoalStep、EgoPERのプロシージャ計画でVLMと構造化プランナーベースラインを破る。
- 参考スコア(独自算出の注目度): 18.644849565753848
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Predicting how a person's first-person view will evolve (what action will follow, what plan completes a task, whether an in-progress shot will score) is fundamentally under-specified: the same context admits many plausible futures, and a model trained to minimize prediction error is forced to hedge or average across them, getting it wrong either way. Two findings shape our approach. First, the future camera trajectory, the path the head carves through space, lets the model commit to one of those futures: it carries the operator's intent in a form fine enough to determine how an action will unfold, substantially outperforming language as a conditioning signal. Second, this same intent makes the trajectory itself partially predictable from the context at hand, enough that trajectory need not be observed at test time to recover most of the gain. We instantiate these findings as TrajPilot, a model that predicts candidate future trajectories from egocentric context and uses them to pilot action prediction in an action-aligned embedding space where language shapes the structure but is never used as a conditioning input. TrajPilot beats VLM and structured-planner baselines on procedural planning across Ego-Exo4D atomic, Ego-Exo4D Keystep, Ego4D GoalStep, and EgoPER, with the trajectory advantage widening with horizon (exactly where prior planners collapse) and holding under RGB-only camera-pose estimation. With the goal masked at inference, the same model performs goal-free anticipation, beating VLM baselines on Ego-Exo4D atomic and extending to EPIC-Kitchens-100 and basketball shot-outcome prediction.
- Abstract(参考訳): 人のファースト・パーソン・ビューがどのように進化するかを予測する(アクション、どのプランがタスクを完了し、プログレッシブ・ショットが得点するかどうか)は、基本的に不明確である。
2つの発見が我々のアプローチを形作っている。
まず、未来のカメラの軌跡、つまり頭部が空間を彫る経路は、オペレーターの意図を十分に詳細に表現し、アクションがどのように展開されるかを判断し、条件付け信号として言語よりもはるかに優れています。
第二に、この同じ意図により、軌道自体が手元にある文脈から部分的に予測可能となり、軌道が利得のほとんどを取り戻すためにテスト時に観測される必要がなくなる。
我々はこれらの知見を,言語が構造を形作るが条件付け入力には使用されないアクション整合型埋め込み空間において,エゴセントリックな文脈から将来の軌道候補を予測するモデルであるTrajPilotとしてインスタンス化する。
TrajPilotはVLMを破り、Ego-Exo4D Atomic、Ego-Exo4D Keystep、Ego4D GoalStep、EgoPERを横断するプロシージャプランニングでベースラインを組む。
推定でマスクされたゴールでは、同じモデルがゴールフリー予測を実行し、Ego-Exo4D原子上でVLMベースラインを破り、EPIC-Kitchens-100まで拡張し、バスケットボールのシュートアウト予測を行う。
関連論文リスト
- Ego-centric Predictive Model Conditioned on Hand Trajectories [52.531681772560724]
自我中心のシナリオでは、次の行動とその視覚的結果の両方を予測することは、人間と物体の相互作用を理解するために不可欠である。
我々は,エゴセントリックなシナリオにおける行動と視覚的未来を共同でモデル化する,統合された2段階予測フレームワークを提案する。
我々のアプローチは、エゴセントリックな人間の活動理解とロボット操作の両方を扱うために設計された最初の統一モデルである。
論文 参考訳(メタデータ) (2025-08-27T13:09:55Z) - Self-Supervised Action-Space Prediction for Automated Driving [0.0]
本稿では,自動走行のための新しい学習型マルチモーダル軌道予測アーキテクチャを提案する。
学習問題を加速度と操舵角の空間に投入することにより、運動論的に実現可能な予測を実現する。
提案手法は,都市交差点とラウンドアバウトを含む実世界のデータセットを用いて評価する。
論文 参考訳(メタデータ) (2021-09-21T08:27:56Z) - Panoptic Segmentation Forecasting [71.75275164959953]
我々の目標は、最近の観測結果から近い将来の予測を行うことです。
この予測能力、すなわち予測能力は、自律的なエージェントの成功に不可欠なものだと考えています。
そこで我々は,2成分モデルを構築した。一方のコンポーネントは,オードメトリーを予測して背景物の力学を学習し,他方のコンポーネントは検出された物の力学を予測する。
論文 参考訳(メタデータ) (2021-04-08T17:59:16Z) - TNT: Target-driveN Trajectory Prediction [76.21200047185494]
我々は移動エージェントのための目標駆動軌道予測フレームワークを開発した。
我々は、車や歩行者の軌道予測をベンチマークする。
私たちはArgoverse Forecasting、InterAction、Stanford Drone、および社内のPedestrian-at-Intersectionデータセットの最先端を達成しています。
論文 参考訳(メタデータ) (2020-08-19T06:52:46Z) - Long-Horizon Visual Planning with Goal-Conditioned Hierarchical
Predictors [124.30562402952319]
未来に予測し、計画する能力は、世界で行動するエージェントにとって基本である。
視覚的予測と計画のための現在の学習手法は、長期的タスクでは失敗する。
本稿では,これらの制約を克服可能な視覚的予測と計画のためのフレームワークを提案する。
論文 参考訳(メタデータ) (2020-06-23T17:58:56Z) - Spatiotemporal Relationship Reasoning for Pedestrian Intent Prediction [57.56466850377598]
視覚データに対する推論は、ロボティクスとビジョンベースのアプリケーションにとって望ましい能力である。
本稿では,歩行者の意図を推論するため,現場の異なる物体間の関係を明らかにするためのグラフ上でのフレームワークを提案する。
歩行者の意図は、通りを横切る、あるいは横断しない将来の行動として定義され、自動運転車にとって非常に重要な情報である。
論文 参考訳(メタデータ) (2020-02-20T18:50:44Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。