論文の概要: WildDet3D: Scaling Promptable 3D Detection in the Wild
- arxiv url: http://arxiv.org/abs/2604.08626v1
- Date: Thu, 09 Apr 2026 16:00:10 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-04-13 17:57:53.502242
- Title: WildDet3D: Scaling Promptable 3D Detection in the Wild
- Title(参考訳): WildDet3D: 野生での3Dのスケーリング
- Authors: Weikai Huang, Jieyu Zhang, Sijun Li, Taoyang Jia, Jiafei Duan, Yunqian Cheng, Jaemin Cho, Mattew Wallingford, Rustin Soraki, Chris Dongjoo Kim, Donovan Clay, Taira Anderson, Winson Han, Ali Farhadi, Bharath Hariharan, Zhongzheng Ren, Ranjay Krishna,
- Abstract要約: テキスト,ポイント,ボックスプロンプトを受信し,推定時に補助的な深度信号を組み込むことができる統合幾何認識アーキテクチャであるWildDet3Dを導入する。
これまでで最大のオープンな3D検出データセットであるWildDet3D-Dataは、既存の2Dアノテーションから候補となる3Dボックスを生成し、人間による検証のみを保持することで構築されている。
- 参考スコア(独自算出の注目度): 67.34808264532982
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Understanding objects in 3D from a single image is a cornerstone of spatial intelligence. A key step toward this goal is monocular 3D object detection--recovering the extent, location, and orientation of objects from an input RGB image. To be practical in the open world, such a detector must generalize beyond closed-set categories, support diverse prompt modalities, and leverage geometric cues when available. Progress is hampered by two bottlenecks: existing methods are designed for a single prompt type and lack a mechanism to incorporate additional geometric cues, and current 3D datasets cover only narrow categories in controlled environments, limiting open-world transfer. In this work we address both gaps. First, we introduce WildDet3D, a unified geometry-aware architecture that natively accepts text, point, and box prompts and can incorporate auxiliary depth signals at inference time. Second, we present WildDet3D-Data, the largest open 3D detection dataset to date, constructed by generating candidate 3D boxes from existing 2D annotations and retaining only human-verified ones, yielding over 1M images across 13.5K categories in diverse real-world scenes. WildDet3D establishes a new state-of-the-art across multiple benchmarks and settings. In the open-world setting, it achieves 22.6/24.8 AP3D on our newly introduced WildDet3D-Bench with text and box prompts. On Omni3D, it reaches 34.2/36.4 AP3D with text and box prompts, respectively. In zero-shot evaluation, it achieves 40.3/48.9 ODS on Argoverse 2 and ScanNet. Notably, incorporating depth cues at inference time yields substantial additional gains (+20.7 AP on average across settings).
- Abstract(参考訳): 単一の画像から3Dで物体を理解することは、空間知性の基盤である。
この目標に向けた重要なステップは、入力されたRGB画像からオブジェクトの広さ、位置、方向を再現するモノクロ3Dオブジェクト検出である。
オープンな世界で実用化するには、そのような検出器は閉集合圏を超えて一般化し、多様な急激なモダリティをサポートし、利用可能であれば幾何学的手がかりを活用する必要がある。
既存のメソッドは1つのプロンプトタイプ用に設計されており、追加の幾何学的キューを組み込むメカニズムが欠如している。
この作業では、両方のギャップに対処します。
まず、テキスト、ポイント、ボックスプロンプトをネイティブに受け付け、推論時に補助的な深度信号を組み込むことができる統合幾何認識アーキテクチャであるWildDet3Dを紹介する。
第2に、既存の2Dアノテーションから候補3Dボックスを生成し、人間認証されたものだけを保持し、13.5Kのカテゴリで100万以上の画像を生成することで、これまでで最大のオープン3D検出データセットであるWildDet3D-Dataを提示する。
WildDet3Dは、複数のベンチマークと設定にまたがって、最先端の新たな状態を確立する。
オープンワールド設定では、新たに導入されたWildDet3D-Benchで22.6/24.8 AP3Dをテキストとボックスプロンプトで達成します。
Omni3Dでは、テキストとボックスプロンプトでそれぞれ34.2/36.4 AP3Dに達する。
ゼロショット評価では、Argoverse 2とScanNetで40.3/48.9 ODSを達成する。
特に、推論時に深さキューを組み込むことで、かなりの利得が得られる(設定全体で平均20.7 AP)。
関連論文リスト
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight [105.9472902251177]
次世代の予測問題として3D検出を行うVLMネイティブレシピを提案する。
このモデルでは, 49.89 AP_3Dの精度を+15.51倍に向上した。
論文 参考訳(メタデータ) (2025-11-25T18:59:45Z) - Towards 3D Objectness Learning in an Open World [19.994404833308092]
我々は,手作りのテキストプロンプトに頼らずに3Dシーン内の物体を検知する,クラス非依存のオープンワールドプロンプトフリー3D検出器OP3Detを提案する。
OP3Detは既存のオープンワールドの3D検出器を最大16.4%超え、クローズドワールドの3D検出器に比べて13.5%改善している。
論文 参考訳(メタデータ) (2025-10-20T16:01:20Z) - 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection [62.57179069154312]
最初のエンドツーエンド3Dモノクロオープンセットオブジェクト検出器(3D-MOOD)を紹介する。
私たちはオープンセットの2D検出を設計した3Dバウンディングボックスヘッドを通して3D空間に持ち上げます。
対象クエリを事前に幾何学的に条件付けし,様々な場面で3次元推定の一般化を克服する。
論文 参考訳(メタデータ) (2025-07-31T13:56:41Z) - Just Add Geometry: Gradient-Free Open-Vocabulary 3D Detection Without Human-in-the-Loop [0.0]
Webスケールのイメージテキストペアでトレーニングされた2次元視覚言語モデルは、リッチなセマンティック理解を示し、オープン語彙検出をサポートする。
我々は,2次元基礎モデルの成熟度とカテゴリの多様性を利用して,人間に注釈を付けた3次元ラベルを使わずに3次元オブジェクト検出を行う。
この結果は,スケーラブルな3D知覚のための2次元基礎モデルの未完成の可能性を強調した。
論文 参考訳(メタデータ) (2025-07-06T15:00:13Z) - General Geometry-aware Weakly Supervised 3D Object Detection [62.26729317523975]
RGB画像と関連する2Dボックスから3Dオブジェクト検出器を学習するための統合フレームワークを開発した。
KITTIとSUN-RGBDデータセットの実験により,本手法は驚くほど高品質な3次元境界ボックスを2次元アノテーションで生成することを示した。
論文 参考訳(メタデータ) (2024-07-18T17:52:08Z) - Uni3D: Exploring Unified 3D Representation at Scale [66.26710717073372]
大規模に統一された3次元表現を探索する3次元基礎モデルであるUni3Dを提案する。
Uni3Dは、事前にトレーニングされた2D ViTのエンドツーエンドを使用して、3Dポイントクラウド機能と画像テキスト整列機能とを一致させる。
強力なUni3D表現は、野生での3D絵画や検索などの応用を可能にする。
論文 参考訳(メタデータ) (2023-10-10T16:49:21Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。