論文の概要: Rarity-Aware Discrete Diffusion with Spatially Consistent Decoding for Photo-Realistic Image Super-Resolution
- arxiv url: http://arxiv.org/abs/2607.17612v2
- Date: Wed, 22 Jul 2026 06:14:49 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-23 16:42:31.589058
- Title: Rarity-Aware Discrete Diffusion with Spatially Consistent Decoding for Photo-Realistic Image Super-Resolution
- Title(参考訳): フォトリアリスティック画像超解像のための空間整合復号を用いた粗さを考慮した離散拡散
- Authors: Ao Li, Yapeng Du, Yi Xin, Lei Zhu, Le Zhang, Guangtao Zhai, Ce Zhu, Xiaohong Liu,
- Abstract要約: DiMOO-SRは光リアル画像超解法のための多モード離散拡散フレームワークである。
広く使われている実世界のSRベンチマークの実験では、DiMOO-SRはいくつかの並列デコードステップで競合する知覚品質を達成する。
- 参考スコア(独自算出の注目度): 80.28905986756685
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Continuous diffusion models have become the dominant paradigm for photo-realistic image Super-Resolution (SR), but they typically formulate reconstruction as continuous signal-level denoising and incorporate semantic priors through external conditioning modules. This makes it less direct to exploit the unified token-based scaling paradigm of modern multimodal models. Autoregressive models provide a more native semantic representation by modeling images as discrete visual tokens, yet their causal decoding is inefficient for high-resolution reconstruction. Discrete diffusion offers a promising middle ground by enabling non-causal, parallel prediction over visual tokens. However, directly adapting discrete diffusion to SR remains non-trivial due to two task-specific challenges: (1) the long-tailed distribution of visual tokens, which under-represents rare but perceptually critical textures; and (2) spatially inconsistent parallel decoding, which may introduce isolated artifacts. To address these issues, we propose DiMOO-SR, a rarity-aware multimodal discrete diffusion framework for photo-realistic image SR. During training, Inverse Frequency Sampling (IFS) prioritizes under-represented but information-rich tokens. During inference, Spatial Consistency Ranking (SCR) refines token confidence using local neighborhood agreement to improve structural coherence. Extensive experiments on widely used real-world SR benchmarks demonstrate that DiMOO-SR achieves competitive perceptual quality with only a few parallel decoding steps, highlighting the potential of discrete diffusion for generative image super-resolution. The code will be released upon publication.
- Abstract(参考訳): 連続拡散モデルは、フォトリアリスティック画像のスーパーリゾリューション(SR)において支配的なパラダイムとなっているが、通常は、連続的な信号レベルの分解として再構成を定式化し、外部条件付きモジュールを通じてセマンティックプリエントを組み込む。
これにより、現代のマルチモーダルモデルのトークンベースの統一スケーリングパラダイムを活用することがより直接的なものになる。
自己回帰モデルは、画像を離散的な視覚トークンとしてモデル化することで、よりネイティブな意味表現を提供するが、それらの因果復号は高解像度再構成では非効率である。
離散拡散は、視覚トークン上の非因果的並列予測を可能にすることによって、有望な中間層を提供する。
しかし、SRに離散拡散を直接適用することは、(1)希少だが知覚的に重要なテクスチャを表現している視覚トークンの長い尾の分布、(2)孤立したアーティファクトを導入できる空間的に一貫性のない並列デコーディングという2つのタスク固有の課題のために、非自明なままである。
これらの問題に対処するため,写真実写画像SRのための多モード離散拡散フレームワークであるDiMOO-SRを提案する。
トレーニング中、IFS(Inverse Frequency Sampling)は、表現されていないが情報に富んだトークンを優先する。
推測中、空間整合性ランキング(SCR)は、局所的な近隣合意を用いてトークンの信頼性を改善し、構造的コヒーレンスを改善する。
広範に使われている実世界のSRベンチマーク実験により、DiMOO-SRは、少数の並列デコードステップで競合する知覚品質を達成し、生成画像超解像における離散拡散の可能性を強調している。
コードは公開時に公開される。
関連論文リスト
- Training-Free Generation of Diverse and High-Fidelity Images via Prompt Semantic Space Optimization [50.5332987313297]
本稿では,トレーニングフリーでモデルに依存しないモジュールであるToken-Prompt Embedding Space Optimization (TPSO)を提案する。
TPSOは、トークン埋め込み空間の未表現領域を探索するために学習可能なパラメータを導入し、学習された分布の強いモードからサンプルを繰り返し生成する傾向を減少させる。
MS-COCOと3つの拡散バックボーンの実験では、TPSOは画像品質を犠牲にすることなく、生成多様性を著しく向上し、ベースライン性能を1.10から4.18ポイントに改善した。
論文 参考訳(メタデータ) (2025-11-25T00:42:09Z) - Visual Autoregressive Modeling for Image Super-Resolution [14.935662351654601]
次世代の予測モデルとして, ISRフレームワークの視覚的自己回帰モデルを提案する。
大規模データを収集し、ロバストな生成先行情報を得るためのトレーニングプロセスを設計する。
論文 参考訳(メタデータ) (2025-01-31T09:53:47Z) - Latent Diffusion, Implicit Amplification: Efficient Continuous-Scale Super-Resolution for Remote Sensing Images [7.920423405957888]
E$2$DiffSRは、最先端のSR手法と比較して、客観的な指標と視覚的品質を達成する。
拡散に基づくSR法の推論時間を非拡散法と同程度のレベルに短縮する。
論文 参考訳(メタデータ) (2024-10-30T09:14:13Z) - Improving Consistency in Diffusion Models for Image Super-Resolution [28.945663118445037]
拡散法における2種類の矛盾を観測する。
セマンティックとトレーニング-推論の組み合わせを扱うために、ConsisSRを導入します。
本手法は,既存拡散モデルにおける最先端性能を示す。
論文 参考訳(メタデータ) (2024-10-17T17:41:52Z) - Binarized Diffusion Model for Image Super-Resolution [61.963833405167875]
超圧縮アルゴリズムであるバイナリ化は、高度な拡散モデル(DM)を効果的に加速する可能性を提供する
既存の二項化法では性能が著しく低下する。
画像SRのための新しいバイナライズ拡散モデルBI-DiffSRを提案する。
論文 参考訳(メタデータ) (2024-06-09T10:30:25Z) - Diffusion-Aided Joint Source Channel Coding For High Realism Wireless Image Transmission [24.372996233209854]
DiffJSCCは条件拡散復調法により高現実性画像を生成する新しいフレームワークである。
768x512ピクセルのコダック画像を3072のシンボルで再現できる。
論文 参考訳(メタデータ) (2024-04-27T00:12:13Z) - Iterative Token Evaluation and Refinement for Real-World
Super-Resolution [77.74289677520508]
実世界の画像超解像(RWSR)は、低品質(LQ)画像が複雑で未同定の劣化を起こすため、長年にわたる問題である。
本稿では,RWSRのための反復的トークン評価・リファインメントフレームワークを提案する。
ITERはGAN(Generative Adversarial Networks)よりも訓練が容易であり,連続拡散モデルよりも効率的であることを示す。
論文 参考訳(メタデータ) (2023-12-09T17:07:32Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。