論文の概要: Monocular Depth Estimation from a Single Image: Progress and Opportunities
- arxiv url: http://arxiv.org/abs/2609.01172v1
- Date: Tue, 01 Sep 2026 12:48:10 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-02 16:31:36.66483
- Title: Monocular Depth Estimation from a Single Image: Progress and Opportunities
- Title(参考訳): 単一画像からの単眼深度推定:進歩と機会
- Authors: Muxin Liu, Xiaoyang Lyu, Yang-Tian Sun, Yi-Hua Huang, Ziyi Yang, Peng Dai, Xiaojuan Qi,
- Abstract要約: この調査は、初期学習に基づく手法からトランスフォーメーション基盤モデルの出現まで、この分野の進化を辿るものである。
本稿では,視覚SLAM,コンテンツ生成,ロボット知覚などのアプリケーションへの深度推定の統合を強調した。
- 参考スコア(独自算出の注目度): 51.389814411876294
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Monocular depth estimation has long stood as a fundamental challenge in computer vision, enabling a wide range of applications including 3D reconstruction, robotics, autonomous driving, and augmented reality. This survey traces the field's evolution from early learning-based methods to the emergence of transformative foundation models. We begin by framing the problem, distinguishing between relative and metric depth estimation, and highlighting the key challenges that have shaped a decade of research. We then present common problem formulations and introduce the most widely used datasets, covering indoor, outdoor, and synthetic data. Following this, we review major advances prior to the foundation model era, distilling core insights from influential methods that contributed to improvements in accuracy, efficiency, and robustness. The survey then turns to the recent surge of foundation-model-based approaches, categorizing them into discriminative and generative paradigms and emphasizing the critical roles of large-scale pretraining (e.g., DINOv3) and synthetic data. We compare representative models using both quantitative benchmarks and qualitative examples, and discuss natural extensions to video-based depth estimation. Further, to illustrate real-world impact, we highlight the integration of depth estimation into applications such as visual SLAM, content generation, and robot perception. Finally, we outline open challenges and promising research directions as the field advances further into the era of foundation models.
- Abstract(参考訳): 単眼深度推定はコンピュータビジョンの基本的な課題として長い間存在しており、3D再構成、ロボティクス、自律運転、拡張現実など幅広い応用を可能にしている。
この調査は、初期学習に基づく手法からトランスフォーメーション基盤モデルの出現まで、この分野の進化を辿るものである。
まず、問題をフレーミングし、相対的な深さとメートル法的な深さの推定を区別し、そして10年にわたる研究を形作る重要な課題を強調します。
次に、一般的な問題定式化を提示し、最も広く使われているデータセットを導入し、屋内、屋外、合成データをカバーした。
これに続いて, 基礎モデル時代以前の大きな進歩を概観し, 精度, 効率, 堅牢性の向上に寄与する影響力ある手法からコアインサイトを抽出した。
この調査は、最近の基礎モデルに基づくアプローチの急増に目を向け、それらを差別的で生成的なパラダイムに分類し、大規模な事前訓練(例えば、DINOv3)と合成データの重要性を強調した。
定量的なベンチマークと定性的な例を用いて代表モデルを比較し、ビデオベース深度推定への自然な拡張について議論する。
さらに、実世界への影響を説明するために、視覚SLAM、コンテンツ生成、ロボット知覚などのアプリケーションへの深度推定の統合を強調した。
最後に、フィールドが基礎モデルの時代にさらに進むにつれて、オープンな課題と有望な研究方向性を概説する。
関連論文リスト
- Towards Depth Foundation Model: Recent Trends in Vision-Based Depth Estimation [96.1872246747684]
深さ推定は3Dコンピュータビジョンの基本課題であり、3D再構成、自由視点レンダリング、ロボティクス、自律運転、AR/VR技術といった応用に不可欠である。
LiDARのようなハードウェアセンサーに依存する従来の方法は、しばしば高コスト、低解像度、環境感度によって制限され、現実のシナリオで適用性を制限する。
ビジョンベースの手法の最近の進歩は有望な代替手段を提供するが、低容量モデルアーキテクチャやドメイン固有の小規模データセットへの依存のため、一般化と安定性の課題に直面している。
論文 参考訳(メタデータ) (2025-07-15T17:59:59Z) - Vision Foundation Models in Remote Sensing: A Survey [6.036426846159163]
ファンデーションモデルは、前例のない精度と効率で幅広いタスクを実行することができる大規模で事前訓練されたAIモデルである。
本調査は, 遠隔センシングにおける基礎モデルの開発と応用を継続するために, 進展のパノラマと将来性のある経路を提供することによって, 研究者や実践者の資源として機能することを目的としている。
論文 参考訳(メタデータ) (2024-08-06T22:39:34Z) - Deep Learning-Based Object Pose Estimation: A Comprehensive Survey [73.74933379151419]
ディープラーニングに基づくオブジェクトポーズ推定の最近の進歩について論じる。
また、複数の入力データモダリティ、出力ポーズの自由度、オブジェクト特性、下流タスクについても調査した。
論文 参考訳(メタデータ) (2024-05-13T14:44:22Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。