論文の概要: Super-Resolution in The Right Latent Space: A Frozen Vision-Foundation Substrate
- arxiv url: http://arxiv.org/abs/2610.04781v1
- Date: Sat, 03 Oct 2026 21:42:53 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-10 02:36:53.40854
- Title: Super-Resolution in The Right Latent Space: A Frozen Vision-Foundation Substrate
- Title(参考訳): 右下肢空間における超解像 : 凍結型視覚創製基板
- Abstract要約: 適切な潜在空間での復元は、SRタスクに不可欠であることを示す。
凍結したDINOv3-Lの23層からなる潜伏空間は、SR作業を容易にするような空間であることを示す。
- 参考スコア(独自算出の注目度): 6.040274700429863
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: In an image latent space, the embeddings of high-resolution, natural, and sharp images form a manifold. Degradation of high-resolution images pushes their embeddings off this manifold. Real-world super-resolution (SR) then becomes the task of mapping the degraded embedding back onto this manifold --- not anywhere on the manifold, but to the point that preserves what the input still carries, both its semantics and pixel details. Every published method implements this mapping in a reconstruction-oriented latent space or pixel space. We claim these spaces are the wrong substrates for SR. Low-resolution and degraded images are embedded far from the manifold, making the mapping difficult and expensive. The lack of semantic information in these substrates also makes it difficult to navigate to the faithful point on the manifold, causing severe hallucination when degradation is heavy. Thus, restoring in a suitable latent space is crucial to the SR task. We show that the latent space of 23 fused layers of a frozen DINOv3-L is one such space that makes the SR task easier. Degraded images are embedded near the manifold. Moreover, this substrate contains a hierarchy of information, from pixel record to degradation robust semantics, guiding the model to find the faithful point on the manifold. On this substrate, a 415M decoder is trained under reconstruction and adversarial objectives to map the degraded embeddings back and decode to pixel space in one pass. The resulting model, RAESR, attains the best fidelity--perception trade-off among state-of-the-art adversarial and diffusion-based restorers on RealSR, DRealSR, LSDIR and DIV2K-Val, at 37 ms per 512 by 512 image on a single H20 GPU. Swapping the substrate for a VAE latent under an identical recipe loses on every metric.
- Abstract(参考訳): 画像潜在空間において、高分解能、自然、および鋭い画像の埋め込みは多様体を形成する。
高解像度画像の劣化は、それらの埋め込みをこの多様体から押し出す。
実世界の超解像(SR)は、分解された埋め込みをこの多様体にマッピングするタスクとなる。
公開されたすべての方法は、このマッピングを再構成指向の潜在空間またはピクセル空間で実装する。
これらの空間は SR の間違った基質であると主張する。
低解像度で劣化した画像は多様体から遠く離れており、マッピングが困難でコストがかかる。
これらの基質に意味情報が欠如しているため、多様体上の忠実な点へのナビゲートも困難であり、劣化が重ければ深刻な幻覚を引き起こす。
したがって、適切な潜在空間での復元は、SRタスクに不可欠である。
凍結したDINOv3-Lの23層からなる潜伏空間は、SR作業を容易にするような空間であることを示す。
劣化した画像は多様体の近くに埋め込まれている。
さらに、この基板は、ピクセルレコードからロバストなセマンティクスの分解に至るまでの情報階層を含んでおり、多様体上の忠実な点を見つけるためにモデルを導く。
この基板上では、415Mデコーダを再構成および対向目的に訓練し、劣化した埋め込みを後方にマッピングし、1パスでピクセル空間にデコーダをデコードする。
結果として得られたRAESRは、RealSR、DRealSR、LSDIR、DIV2K-Valの最先端の敵対者および拡散ベースの復元者の間で、単一のH20 GPU上の512イメージで37ms/512で、最高の忠実度-知覚トレードオフを達成する。
同一のレシピの下でVAE潜伏剤の基質をスワップすると、すべてのメートル法で失われる。
関連論文リスト
- PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation [105.86385617182879]
最先端のシングルイメージの3D再構成手法は複雑なハイブリッドアーキテクチャと損失関数に依存していることが多い。
このようなアーキテクチャ上のオーバーヘッドと複雑な損失の定式化は不要であることを示す。
我々は,平易なViT上に構築された最小限の画素空間拡散変換器を導入し,生の3Dポイントマップパッチを直接操作する。
論文 参考訳(メタデータ) (2026-07-02T17:59:56Z) - Super-Resolution through StyleGAN Regularized Latent Search: A
Realism-Fidelity Trade-off [3.212648064850423]
本稿では,高分解能(HR)画像を低分解能(LR)画像から構築する問題に対処する。
最近の教師なしアプローチでは、HR画像上で事前訓練されたStyleGANの潜伏空間を探索し、入力LR画像に最もダウンスケールした画像を求める。
我々は、潜在空間における探索を制約する新しい正規化器を導入し、逆符号が元の画像多様体に存在することを保証する。
論文 参考訳(メタデータ) (2023-11-28T16:27:24Z) - Space Debris: Are Deep Learning-based Image Enhancements part of the
Solution? [9.117415383776695]
現在地球を周回している宇宙ゴミの量は、加速ペースで持続不可能なレベルに達している。
軌道定義された、登録された宇宙船と、ローグ/非活動的な宇宙物体の検知、追跡、識別、識別は、資産保護に不可欠である。
本研究の主な目的は、可視光スペクトルの単眼カメラで捉えた際の限界や画像アーチファクトを克服するために、ディープニューラルネットワーク(DNN)ソリューションの有効性を検討することである。
論文 参考訳(メタデータ) (2023-08-01T09:38:41Z) - Symmetric Uncertainty-Aware Feature Transmission for Depth
Super-Resolution [52.582632746409665]
カラー誘導DSRのためのSymmetric Uncertainty-aware Feature Transmission (SUFT)を提案する。
本手法は最先端の手法と比較して優れた性能を実現する。
論文 参考訳(メタデータ) (2023-06-01T06:35:59Z) - Towards Lightweight Super-Resolution with Dual Regression Learning [58.98801753555746]
深層ニューラルネットワークは、画像超解像(SR)タスクにおいて顕著な性能を示した。
SR問題は通常不適切な問題であり、既存の手法にはいくつかの制限がある。
本稿では、SRマッピングの可能な空間を削減するために、二重回帰学習方式を提案する。
論文 参考訳(メタデータ) (2022-07-16T12:46:10Z) - High-resolution Depth Maps Imaging via Attention-based Hierarchical
Multi-modal Fusion [84.24973877109181]
誘導DSRのための新しい注意に基づく階層型マルチモーダル融合ネットワークを提案する。
本手法は,再現精度,動作速度,メモリ効率の点で最先端手法よりも優れていることを示す。
論文 参考訳(メタデータ) (2021-04-04T03:28:33Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。