FuguReport

AerialMetric: Benchmarking and Adapting UAV Monocular Metric Depth Estimation in the Real World

Authors Zhongqiang Song, Guanying Chen, Yuqi Zhang, Yin Zou, Chuanyu Fu, Zhiyuan Yuan, Chuan Huang, Shuguang Cui, Xiaochun Cao
Affiliations FNii Shenzhen / SIAS USTC / SSE CUHKSZ / Sun Yat-Sen University
Categories Evaluation / Benchmarking / Monocular depth estimation benchmark, Application / UAV Perception / Depth estimation for aerial UAVs, Method / Domain Adaptation / Adapting metric depth estimation to real world
License CC BY 4.0

Abstract Overview

This paper studies monocular metric depth estimation for UAV aerial imagery, where models trained on ground-level data show substantial domain gaps. To support evaluation and adaptation, the authors introduce AerialMetric, a benchmark composed of four subsets: real-world oblique photogrammetry data, a controlled decoupled UAV test set, photorealistic synthetic data, and in-the-wild internet drone videos with pseudo-metric labels. The dataset contains 52K real-world and 16K synthetic image-depth pairs with metric or pseudo-metric supervision, and is designed to analyze the effects of viewpoint, altitude, and camera field of view. Using this benchmark, the paper evaluates several existing metric depth estimators and fine-tunes a representative model to assess how aerial-specific training changes performance.

Novelty

The main novelty is the construction of a UAV-focused metric depth benchmark that combines large-scale real and synthetic data with two distinctive evaluation settings: a variable-decoupled acquisition protocol and an in-the-wild pseudo-metric test set. This setup enables systematic analysis of aerial imaging factors that are usually entangled or missing in prior depth benchmarks.

Results

The experiments show that existing zero-shot metric depth models transfer poorly to aerial imagery, often with very high AbsRel and near-zero δ1 on AerialMetric. After fine-tuning MoGe2 on AerialMetric, performance improves substantially across aerial benchmarks; for example, on AerialMetric-Oblique-City without ground-truth intrinsics, δ1 rises from 5.1 to 89.3, and on AerialMetric-Wild (0–400 m), AbsRel drops from 55.26 to 21.34 while δ1 increases from 9.5 to 53.7. The adapted model also remains competitive on several ground-domain benchmarks, indicating limited loss of cross-domain generalization.

Key Points

  1. AerialMetric combines four complementary subsets to cover curated real UAV imagery, controlled decoupled captures, synthetic training data, and in-the-wild aerial videos.
  2. The benchmark exposes a severe domain gap between ground-trained monocular metric depth models and aerial UAV viewpoints, including sensitivity to altitude, pitch, and field of view.
  3. Fine-tuning a representative model on AerialMetric yields large gains on aerial depth estimation while preserving competitive performance on multiple ground-domain datasets.

References

This page was created using generative AI such as GPT-5, Claude Opus 4, Gemini 3, Gemini 3.1 Flash Image, and their higher-end successor versions. No guarantee can be made regarding its contents.