論文の概要: Human-in-the-Loop Geospatial Annotation for Rapid Dataset Construction in Field-Deployed UAV Systems
- arxiv url: http://arxiv.org/abs/2609.28767v3
- Date: Tue, 29 Sep 2026 15:25:56 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-30 21:28:46.531245
- Title: Human-in-the-Loop Geospatial Annotation for Rapid Dataset Construction in Field-Deployed UAV Systems
- Title(参考訳): フィールド展開型UAVシステムにおける高速データセット構築のための人対地空間アノテーション
- Abstract要約: 本稿では、エキスパートアノテーションを画像からフィールドにシフトするBirdsEyeを紹介する。
カメラ目的の不確実性から画素不確実性への一階写像を導出し,モンテカルロシミュレーションに対して検証する。
システム投影精度は10~20mのAGL高度でサブディシメータ(30ピクセル未満)であることを示す。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by-nc-sa/4.0/
- Abstract: Real-world perception systems must adapt to changing environments, but manual image annotation cannot scale to field data volumes. We present BirdsEye, which shifts expert annotation from images to the field: an operator records target locations in world coordinates using RTK positioning and calibrated projective geometry propagates each observation to all frames where the target is visible. To quantify how well physical annotations align with image observations, we derive a first-order mapping from camera-pose uncertainty to pixel uncertainty and validate it against Monte Carlo simulation. This mapping is linear in the six per-axis pose variances, so it inverts into a sensor design tool: we give a sufficient condition converting an annotation tolerance into a convex set of admissible pose-noise budgets, a closed-form largest admissible scaling of a deployed sensor suite, and a unique per-axis pose specification under an equal-budget-share allocation. We also analyze the planar-surface approximation underlying the projection, which holds up to 10 degrees of terrain slope. By direct measurement, we show that system projection accuracy is sub-decimeter (sub-30 pixel) at AGL altitudes of 10-20m under conditions excluding sustained yawing. During an in-field case study across three agricultural sites, two field workers produced 12,524 annotated frames carrying 55,600 labels in roughly 12 hours (25.5x per-worker rate increase over manual labeling). Detectors trained on imagery collected by this workflow recovered 56-89% of in-view surveyed targets at a geographically distinct farm, at pre-registered operating points; human review of the leading configuration estimates detection precision at 83-87%, spanning three tie-break conventions for clusters carrying contradictory human verdicts.
- Abstract(参考訳): 実世界の知覚システムは環境の変化に適応する必要があるが、手動画像アノテーションはデータボリュームを拡大できない。
演算子はRTK測位と校正された射影幾何を用いて世界座標の目標位置を記録し、それぞれの観測を目標が見える全てのフレームに伝搬する。
物理的アノテーションが画像観察とどのように一致しているかを定量化するために、カメラの目的の不確実性から画素の不確実性への一階写像を導出し、モンテカルロシミュレーションに対して検証する。
このマッピングは6軸のポーズのばらつきに線形であり、センサ設計ツールに逆転する: 許容されたポーズの予算の凸集合にアノテーションのトレランスを変換する十分な条件、配置されたセンサースイートの最大許容スケーリングをクローズドな形式とし、均等な予算シェアアロケーションの下でアクセント毎のポーズ仕様を与える。
また,地形の傾斜度を最大10度に抑えるプロジェクションの基礎となる平面面近似を解析した。
直接測定により, 連続ヨーイングを除く条件下でのAGL高度10~20mでは, システム投影精度がサブディシメータ(30ピクセル以下)であることが確認された。
3つの農地の現場調査において、2人の現場労働者が約12時間で55,600個のラベルを積んだ12,524個の注釈付きフレームを生産した(作業者1人当たり25.5倍)。
このワークフローによって収集された画像に基づいてトレーニングされた検出器は、地理的に異なる農場で調査対象の56~89%を、事前登録された運用ポイントで回収した。
関連論文リスト
- FoundationGeo: Learning Spatial Pixel-Wise Fields for Monocular Metric Geometry [52.232826445242644]
FoundationGeoは空間キャリブレーションと原則データ設計による相対的および計量的予測を橋渡しする。
ステージ1は DINOv3 で初期化することで高忠実なアフィン不変幾何モデルを学ぶ。
ステージ2は、メートル法推定のためのピクセルワイドキャリブレーションフィールドを導入することで、グローバルなスケーリングを越えている。
論文 参考訳(メタデータ) (2026-07-13T14:10:01Z) - Unifying UAV Cross-View Geo-Localization via 3D Geometric Perception [51.687842983240564]
無人航空機(UAV)のクロスビューな地上局地化は、斜めのUAV画像と衛星地図との厳密な幾何学的相違により、いまだに困難である。
本稿では,3次元シーン形状を明示的にモデル化し,粗い位置認識ときめ細かなポーズ推定を統一する,幾何認識型UAV測位フレームワークを提案する。
提案手法は, 最先端のベースラインを著しく上回り, ロバストメータレベルのローカライゼーション精度を実現し, 複雑な都市環境における一般化を向上する。
論文 参考訳(メタデータ) (2026-04-02T08:08:41Z) - Loc$^2$: Interpretable Cross-View Localization via Depth-Lifted Local Feature Matching [80.57282092735991]
本稿では,高精度かつ解釈可能なクロスビューローカライズ手法を提案する。
地上画像の3自由度(DoF)のポーズを、その局所的な特徴と基準空中画像とをマッチングすることによって推定する。
実験では、クロスエリアテストや未知の向きといった挑戦的なシナリオにおいて、最先端の精度を示す。
論文 参考訳(メタデータ) (2025-09-11T18:52:16Z) - Convolutional Cross-View Pose Estimation [9.599356978682108]
クロスビューポーズ推定のための新しいエンドツーエンド手法を提案する。
提案手法は,VIGORおよびKITTIデータセット上で検証される。
オックスフォード・ロボットカーのデータセットでは,エゴ車両の姿勢を時間とともに確実に推定することができる。
論文 参考訳(メタデータ) (2023-03-09T13:52:28Z) - A Large Scale Homography Benchmark [52.55694707744518]
1DSfMデータセットから10万枚の画像から約1000個の平面が観測された3D, Pi3Dの平面の大規模データセットを示す。
また,Pi3Dを利用した大規模ホモグラフィ推定ベンチマークであるHEBを提案する。
論文 参考訳(メタデータ) (2023-02-20T14:18:09Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。