Fugu-MT 論文翻訳(概要): HV-BEV: Decoupling Horizontal and Vertical Feature Sampling for Multi-View 3D Object Detection

論文の概要: HV-BEV: Decoupling Horizontal and Vertical Feature Sampling for Multi-View 3D Object Detection

arxiv url: http://arxiv.org/abs/2412.18884v2
Date: Mon, 30 Dec 2024 13:49:45 GMT
ステータス: 翻訳完了
システム内更新日: 2024-12-31 14:20:38.411952
Title: HV-BEV: Decoupling Horizontal and Vertical Feature Sampling for Multi-View 3D Object Detection
Title（参考訳）: HV-BEV:多視点3次元物体検出のための水平・垂直特徴サンプリングの分離
Authors: Di Wu, Feng Yang, Benlian Xu, Pan Liao, Wenhui Zhao, Dingwen Zhang,
Abstract要約: HV-BEVは、BEVグリッドクエリのパラダイムにおける特徴サンプリングを水平特徴集約と垂直適応高さ対応基準点サンプリングに分離する新しいアプローチである。我々の最高のパフォーマンスモデルは、nuScenesテストセットで50.5%のmAPと59.8%のNDSを達成する。
参考スコア（独自算出の注目度）: 34.72603963887331
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Abstract: The application of vision-based multi-view environmental perception system has been increasingly recognized in autonomous driving technology, especially the BEV-based models. Current state-of-the-art solutions primarily encode image features from each camera view into the BEV space through explicit or implicit depth prediction. However, these methods often focus on improving the accuracy of projecting 2D features into corresponding depth regions, while overlooking the highly structured information of real-world objects and the varying height distributions of objects across different scenes. In this work, we propose HV-BEV, a novel approach that decouples feature sampling in the BEV grid queries paradigm into horizontal feature aggregation and vertical adaptive height-aware reference point sampling, aiming to improve both the aggregation of objects' complete information and generalization to diverse road environments. Specifically, we construct a learnable graph structure in the horizontal plane aligned with the ground for 3D reference points, reinforcing the association of the same instance across different BEV grids, especially when the instance spans multiple image views around the vehicle. Additionally, instead of relying on uniform sampling within a fixed height range, we introduce a height-aware module that incorporates historical information, enabling the reference points to adaptively focus on the varying heights at which objects appear in different scenes. Extensive experiments validate the effectiveness of our proposed method, demonstrating its superior performance over the baseline across the nuScenes dataset. Moreover, our best-performing model achieves a remarkable 50.5% mAP and 59.8% NDS on the nuScenes testing set.
Abstract（参考訳）: 視覚に基づくマルチビュー環境認識システムの応用は、自律運転技術、特にBEVベースのモデルにおいてますます認識されている。現在の最先端ソリューションは主に、暗黙の深度予測を通じて、各カメラビューからの画像をBEV空間にエンコードする。しかし、これらの手法は、実世界の物体の高度情報や異なる場面における物体の高さ分布を網羅しながら、2次元特徴を対応する深度領域に投影する精度の向上に重点を置いていることが多い。本研究では,BEVグリッドクエリのパラダイムにおける特徴サンプリングを水平的特徴集約と垂直適応型高さ対応基準点サンプリングに分離する新しい手法であるHV-BEVを提案する。具体的には,3次元基準点を接地した水平面に学習可能なグラフ構造を構築し,特に車両周辺の複数の画像ビューにまたがる場合において,異なるBEVグリッドにまたがる同一のインスタンスの関連を補強する。また,固定高さ範囲内における一様サンプリングに頼る代わりに,歴史的情報を組み込んだ高さ認識モジュールを導入し,参照ポイントが異なるシーンに物体が現れる様々な高さに適応的に焦点を合わせることができるようにした。大規模な実験により提案手法の有効性が検証され, nuScenesデータセットのベースラインよりも優れた性能を示した。さらに、我々の最高のパフォーマンスモデルは、nuScenesテストセット上で50.5%のmAPと59.8%のNDSを達成する。

関連論文リスト

RaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera Fusion [58.77329237533034]
本稿では3次元物体検出の精度を高めるために,Raar-Camera fusion transformer (RaCFormer)を提案する。 RaCFormerは、nuScenesデータセット上で64.9% mAPと70.2%の優れた結果を得る。
論文参考訳（メタデータ） (2024-12-17T09:47:48Z)
HeightLane: BEV Heightmap guided 3D Lane Detection [6.940660861207046]
単分子画像からの正確な3次元車線検出は、深さのあいまいさと不完全な地盤モデリングによる重要な課題を示す。本研究は,マルチスロープ仮定に基づいてアンカーを作成することにより,単眼画像から高さマップを予測する革新的な手法であるHeightLaneを紹介する。 HeightLaneは、Fスコアの観点から最先端のパフォーマンスを実現し、現実世界のアプリケーションにおけるその可能性を強調している。
論文参考訳（メタデータ） (2024-08-15T17:14:57Z)
Towards Unified 3D Object Detection via Algorithm and Data Unification [70.27631528933482]
我々は、最初の統一型マルチモーダル3Dオブジェクト検出ベンチマークMM-Omni3Dを構築し、上記のモノクロ検出器をマルチモーダルバージョンに拡張する。設計した単分子・多モード検出器をそれぞれUniMODEとMM-UniMODEと命名した。
論文参考訳（メタデータ） (2024-02-28T18:59:31Z)
Instance-aware Multi-Camera 3D Object Detection with Structural Priors Mining and Self-Boosting Learning [93.71280187657831]
カメラによる鳥眼視(BEV)知覚パラダイムは、自律運転分野において大きな進歩を遂げている。画像平面のインスタンス認識をBEV検出器内の深度推定プロセスに統合するIA-BEVを提案する。
論文参考訳（メタデータ） (2023-12-13T09:24:42Z)
Towards Generalizable Multi-Camera 3D Object Detection via Perspective Debiasing [28.874014617259935]
マルチカメラ3Dオブジェクト検出(MC3D-Det)は,鳥眼ビュー(BEV)の出現によって注目されている。本研究では,3次元検出と2次元カメラ平面との整合性を両立させ,一貫した高精度な検出を実現する手法を提案する。
論文参考訳（メタデータ） (2023-10-17T15:31:28Z)
CoBEV: Elevating Roadside 3D Object Detection with Depth and Height Complementarity [34.025530326420146]
我々は、新しいエンドツーエンドのモノクロ3Dオブジェクト検出フレームワークであるComplementary-BEVを開発した。道路カメラを用いたDAIR-V2X-IとRope3Dの公開3次元検出ベンチマークについて広範な実験を行った。カメラモデルのAPスコアが初めてDAIR-V2X-Iで80%に達する。
論文参考訳（メタデータ） (2023-10-04T13:38:53Z)
OCBEV: Object-Centric BEV Transformer for Multi-View 3D Object Detection [29.530177591608297]
マルチビュー3Dオブジェクト検出は、高い有効性と低コストのため、自動運転において人気を博している。現在の最先端検出器のほとんどは、クエリベースのバードアイビュー(BEV)パラダイムに従っている。本稿では,移動対象の時間的・空間的手がかりをより効率的に彫ることができるOCBEVを提案する。
論文参考訳（メタデータ） (2023-06-02T17:59:48Z)
BEV-IO: Enhancing Bird's-Eye-View 3D Detection with Instance Occupancy [58.92659367605442]
我々は,BEV表現をインスタンス占有情報で拡張する新しい3次元検出パラダイムであるBEV-IOを提案する。 BEV-IOは、パラメータや計算オーバーヘッドの無視できる増加しか加えず、最先端の手法よりも優れていることを示す。
論文参考訳（メタデータ） (2023-05-26T11:16:12Z)
Geometric-aware Pretraining for Vision-centric 3D Object Detection [77.7979088689944]
GAPretrainと呼ばれる新しい幾何学的事前学習フレームワークを提案する。 GAPretrainは、複数の最先端検出器に柔軟に適用可能なプラグアンドプレイソリューションとして機能する。 BEVFormer法を用いて, nuScenes val の 46.2 mAP と 55.5 NDS を実現し, それぞれ 2.7 と 2.1 点を得た。
論文参考訳（メタデータ） (2023-04-06T14:33:05Z)
OA-BEV: Bringing Object Awareness to Bird's-Eye-View Representation for Multi-Camera 3D Object Detection [78.38062015443195]
OA-BEVは、BEVベースの3Dオブジェクト検出フレームワークにプラグインできるネットワークである。提案手法は,BEV ベースラインに対する平均精度と nuScenes 検出スコアの両面で一貫した改善を実現する。
論文参考訳（メタデータ） (2023-01-13T06:02:31Z)

関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。

指定された論文の情報です。
本サイトの運営者は本サイト（すべての情報・翻訳含む）の品質を保証せず、本サイト（すべての情報・翻訳含む）を使用して発生したあらゆる結果について一切の責任を負いません。