論文の概要: IndoorBEV: A Lightweight Real-Time LiDAR BEV Perception System for Indoor Mobile Robots
- arxiv url: http://arxiv.org/abs/2610.00355v1
- Date: Wed, 30 Sep 2026 00:37:20 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-03 01:19:23.618301
- Title: IndoorBEV: A Lightweight Real-Time LiDAR BEV Perception System for Indoor Mobile Robots
- Title(参考訳): 屋内移動ロボットのための軽量リアルタイムLiDARBEV知覚システムIndoorBEV
- Abstract要約: 移動ロボットは、厳しいレイテンシとメモリ制約の下で、散らかった3次元環境を理解する必要がある。
本稿では,このトレードオフを軽減する軽量LiDAR認識フレームワークであるIndoorBEVを紹介する。
IndoorBEVは0.6Mパラメータしか持たず、2.3MBのモデルストレージを必要とする。
- 参考スコア(独自算出の注目度): 0.0
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Efficient indoor LiDAR perception is challenging because mobile robots must understand cluttered three-dimensional environments under strict latency and memory constraints. Existing point-based and voxel-based methods often incur substantial computational overhead, whereas conventional bird's-eye-view (BEV) representations improve efficiency at the cost of discarding vertical geometric information. We present IndoorBEV, a lightweight LiDAR perception framework that mitigates this tradeoff through a height-aware BEV representation and geometry-conditioned feature fusion. IndoorBEV summarizes the vertical point distribution in each BEV cell using statistical height features and multi-frequency height encoding, allowing informative three-dimensional cues to be processed efficiently by two-dimensional convolutions. A lightweight encoder then integrates complementary geometric features with multi-scale local representations and compact global scene context. Decoupled dense prediction heads jointly produce semantic BEV maps and oriented object bounding boxes. IndoorBEV contains only 0.6M parameters and requires 2.3 MB of model storage. On an NVIDIA AGX Orin, it uses 21.52 MB of GPU memory per inference and achieves a mean latency of 169.6 ms under a 200 ms perception deadline, with a deadline miss ratio of 1.8\%. Evaluations on simulated scenes, real-world robot scans, and an open-source indoor point-cloud dataset demonstrate a favorable tradeoff among perception accuracy, latency, and memory consumption. These results indicate that explicitly encoding vertical geometry within a compact BEV representation provides an effective approach to resource-efficient indoor LiDAR perception.
- Abstract(参考訳): 移動ロボットは、厳密なレイテンシとメモリ制約の下で、散らばった3次元環境を理解しなければならないため、効率的な屋内LiDAR認識は困難である。
既存の点ベースおよびボクセルベースの手法は、しばしばかなりの計算オーバーヘッドを発生させるが、従来の鳥眼ビュー(BEV)表現は垂直幾何情報を捨てるコストで効率を向上させる。
本稿では,このトレードオフを軽減する軽量LiDAR認識フレームワークであるIndoorBEVを提案する。
IndoorBEVは、統計的高さ特徴と多周波高符号化を用いて各BEVセルの垂直点分布を要約し、2次元畳み込みにより情報的3次元キューを効率的に処理する。
軽量エンコーダは、補完幾何学的特徴を多スケール局所表現とコンパクトなグローバルシーンコンテキストと統合する。
分離された密接な予測ヘッドは、セマンティックなBEVマップとオブジェクト指向のオブジェクト境界ボックスを共同で生成する。
IndoorBEVは0.6Mパラメータしか持たず、2.3MBのモデルストレージを必要とする。
NVIDIA AGX Orinでは、推論毎に21.52MBのGPUメモリを使用し、200ミリ秒の認識期限下で平均レイテンシ169.6msを達成する。
シミュレーションシーン、実世界のロボットスキャン、オープンソースの屋内ポイントクラウドデータセットの評価は、認識精度、レイテンシ、メモリ消費のトレードオフとして好ましいことを示している。
これらの結果は、コンパクトなBEV表現内の垂直形状を明示的に符号化することは、資源効率の高い屋内LiDAR知覚への効果的なアプローチを提供することを示している。
関連論文リスト
- Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints [20.26519299938903]
オープンボキャブラリBEVセグメンテーション(OVBS)を導入し、トレーニングセットを超えてカテゴリを認識する。
OVBEVSegはジオメトリを意識したOVBSフレームワークであり、効率的なガウススプラッティング(GS)ベースのアンプロジェクションを強化する。
nuScenesデータセットでは、OVBEVSegは最先端のパフォーマンスを達成し、未確認のカテゴリで15.3 mIoUのクローズドセットメソッドよりも優れています。
論文 参考訳(メタデータ) (2026-06-23T09:43:12Z) - Lightweight Spatial Embedding for Vision-based 3D Occupancy Prediction [37.8001844396061]
LightOccは、軽量空間埋め込みを利用する革新的な3D占有予測フレームワークである。
LightOccはベースラインの予測精度を大幅に向上させ、Occ3D-nuScenesベンチマークで最先端のパフォーマンスを達成する。
論文 参考訳(メタデータ) (2024-12-08T15:49:35Z) - U-BEV: Height-aware Bird's-Eye-View Segmentation and Neural Map-based Relocalization [81.76044207714637]
GPS受信が不十分な場合やセンサベースのローカライゼーションが失敗する場合、インテリジェントな車両には再ローカライゼーションが不可欠である。
Bird's-Eye-View (BEV)セグメンテーションの最近の進歩は、局所的な景観の正確な推定を可能にする。
本稿では,U-NetにインスパイアされたアーキテクチャであるU-BEVについて述べる。
論文 参考訳(メタデータ) (2023-10-20T18:57:38Z) - BEV-IO: Enhancing Bird's-Eye-View 3D Detection with Instance Occupancy [58.92659367605442]
我々は,BEV表現をインスタンス占有情報で拡張する新しい3次元検出パラダイムであるBEV-IOを提案する。
BEV-IOは、パラメータや計算オーバーヘッドの無視できる増加しか加えず、最先端の手法よりも優れていることを示す。
論文 参考訳(メタデータ) (2023-05-26T11:16:12Z) - BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird's-Eye View Representation [105.96557764248846]
本稿では,汎用マルチタスクマルチセンサ融合フレームワークであるBEVFusionを紹介する。
共有鳥眼ビュー表示空間におけるマルチモーダル特徴を統一する。
3Dオブジェクト検出では1.3%高いmAPとNDS、BEVマップのセグメンテーションでは13.6%高いmIoU、コストは1.9倍である。
論文 参考訳(メタデータ) (2022-05-26T17:59:35Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。