論文の概要: Rethinking Monocular Depth Embedding for Generalized Stereo Matching
- arxiv url: http://arxiv.org/abs/2607.09284v1
- Date: Fri, 10 Jul 2026 10:51:06 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-13 14:47:12.832212
- Title: Rethinking Monocular Depth Embedding for Generalized Stereo Matching
- Title(参考訳): 一般化ステレオマッチングのための単眼深度埋め込みの再考
- Authors: Libo Lin, Shuangli Du, Minghua Zhao, Zhenzhen You, Shun Lv, Yiguang Liu,
- Abstract要約: 単分子的手法は、リッチな文脈的先行を捉えるが、幾何学的精度は欠くが、ステレオ的手法は幾何学的に正確であり、テクスチャのない領域で苦戦している。
いくつかのアプローチは、ステレオ情報と単分子深度を整合させることにより、ステレオマッチング(SM)の一般化を促進するために、それらの強みを組み合わせる。
本稿では,単分子深度誤差に対する耐性を向上させるために,単分子深度からのハード制約ではなくソフト制約を用いて,単分子深度埋め込みを再考する。
- 参考スコア(独自算出の注目度): 7.3558734052495245
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Generally, monocular methods capture rich contextual priors but lack geometric precision, whereas stereo methods are geometrically accurate yet struggle in textureless and occluded regions. Several approaches attempt to combine their strengths to enhance the generalization of stereo matching (SM) by aligning monocular depth with stereo information. However, establishing a stable and generalizable alignment is challenging, and unreliable monocular cues can substantially degrade performance. This paper rethinks monocular depth embedding. First, to prevent shortcut learning, we reduce branch coupling instead of expanding network width. Second, we construct soft constraints instead of hard ones from monocular depth to improve tolerance to monocular depth errors. Based on the principles, we integrate monocular information into both feature extraction and GRU iterations. Specifically, the monocular depth map is fused with the RGB image to sharpen depth boundary perception and suppress matching ambiguities. The fused image is then used for feature extraction, allowing the contextual features to encode global geometric information. Furthermore, the monocular depth gradient feature is employed to guide disparity updates, helping to escape local oscillations. Finally, to address the boundary blurring of supervised disparity caused by data augmentation, we propose an edge confidence estimation method and an edge-aware loss function. Our method achieves state-of-the-art (SOTA) performance on multiple standard benchmarks, demonstrating excellent generalization while improving accuracy. The code is available at https://github.com/linliboabc-maker/stereo-matching-digital.
- Abstract(参考訳): 一般に、単分子的手法はリッチな文脈的先行を捉えるが、幾何学的精度は欠くが、ステレオ的手法は幾何学的に正確であり、テクスチャのない領域では困難である。
いくつかのアプローチは、ステレオ情報と単分子深度を整合させることにより、ステレオマッチング(SM)の一般化を促進するために、それらの強みを組み合わせる。
しかし、安定かつ一般化可能なアライメントを確立することは困難であり、信頼性の低い単分子キューは性能を著しく低下させる可能性がある。
本稿では, 単分子深度埋め込みについて再考する。
まず、ショートカット学習を防止するために、ネットワーク幅を広げる代わりに分岐結合を減らす。
第二に、単分子深度誤差に対する耐性を向上させるために、単分子深度からのハード制約の代わりにソフト制約を構築する。
原理に基づいて、単分子情報を特徴抽出とGRU反復の両方に統合する。
具体的には、単眼深度マップをRGB画像と融合させ、深度境界知覚を鋭くし、一致したあいまいさを抑制する。
融合した画像は特徴抽出に使用されるので、文脈的特徴はグローバルな幾何学的情報をエンコードすることができる。
さらに、単分子深度勾配特徴は、不均一な更新を誘導し、局所的な振動から逃れるのに役立つ。
最後に,データ拡張による教師付き不一致の境界のぼかしに対処するため,エッジ信頼度推定法とエッジ認識損失関数を提案する。
提案手法は,複数の標準ベンチマーク上でのSOTA(State-of-the-art)性能を実現し,精度を向上しながら優れた一般化を実現する。
コードはhttps://github.com/linliboabc-maker/stereo-matching-digitalで公開されている。
関連論文リスト
- MDE-VIO: Enhancing Visual-Inertial Odometry Using Learned Depth Priors [8.2208199207543]
本稿では,アフィン不変深度一貫性と対方向順序制約を強制する新しいフレームワークを提案する。
このアプローチは、計量スケールを頑健に回復しながら、エッジデバイスの計算限界に厳密に固執する。
論文 参考訳(メタデータ) (2026-02-11T19:53:06Z) - BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent Alignment [31.118114556998048]
モノラルおよびステレオのアプローチを3次元推定にブリッジする統合フレームワークを導入する。
新しいクロスアテンタティブアライメント機構は、ステレオ仮説表現とモノクロコンテキストキューを動的に同期させる。
我々のアプローチは、モダリティ固有の制限を超越した堅牢な3D知覚を可能にする。
論文 参考訳(メタデータ) (2025-08-06T16:31:22Z) - Diving into the Fusion of Monocular Priors for Generalized Stereo Matching [27.15757281613792]
近年,視覚基礎モデル (VFM) に先立って, 偏りのない単分子を応用して, 不測領域の一般化を向上することで, ステレオマッチングが進展している。
本稿では,深度マップを二項相対形式に変換する融合を導くための二項局所順序付けマップを提案する。
また、画素単位の線形回帰モジュールがそれらをグローバルかつ適応的に整列できるような登録問題として、単分子深度を不均質に最終的に直接融合させることを定式化する。
論文 参考訳(メタデータ) (2025-05-20T14:27:45Z) - Marigold-DC: Zero-Shot Monocular Depth Completion with Guided Diffusion [57.08169927189237]
奥行き完了のための既存の手法は、厳密に制約された設定で動作する。
単眼深度推定の進歩に触発されて,画像条件の深度マップ生成として深度補完を再構成した。
Marigold-DCは、単分子深度推定のための事前訓練された潜伏拡散モデルを構築し、試験時間ガイダンスとして深度観測を注入する。
論文 参考訳(メタデータ) (2024-12-18T00:06:41Z) - AugUndo: Scaling Up Augmentations for Monocular Depth Completion and Estimation [51.143540967290114]
本研究では,教師なし深度計算と推定のために,従来不可能であった幾何拡張の幅広い範囲をアンロックする手法を提案する。
これは、出力深さの座標への幾何変換を反転、あるいはアンドウイング(undo''-ing)し、深度マップを元の参照フレームに戻すことで達成される。
論文 参考訳(メタデータ) (2023-10-15T05:15:45Z) - DevNet: Self-supervised Monocular Depth Learning via Density Volume
Construction [51.96971077984869]
単眼画像からの自己教師付き深度学習は、通常、時間的に隣接する画像フレーム間の2Dピクセル単位の光度関係に依存する。
本研究は, 自己教師型単眼深度学習フレームワークであるDevNetを提案する。
論文 参考訳(メタデータ) (2022-09-14T00:08:44Z) - Occlusion-Aware Depth Estimation with Adaptive Normal Constraints [85.44842683936471]
カラービデオから多フレーム深度を推定する新しい学習手法を提案する。
本手法は深度推定精度において最先端の手法より優れる。
論文 参考訳(メタデータ) (2020-04-02T07:10:45Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。