論文の概要: Double Descent and Malign Overfitting in Diffusion Models
- arxiv url: http://arxiv.org/abs/2609.26392v1
- Date: Tue, 22 Sep 2026 13:30:48 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-23 18:04:04.391939
- Title: Double Descent and Malign Overfitting in Diffusion Models
- Title(参考訳): 拡散モデルにおける二重発色と悪性オーバーフィッティング
- Abstract要約: 拡散モデルの過度な適合は破滅的であり、モデルを体制へと駆り立てる。
トレーニングサンプルあたりの雑音実現の固定数$m$では、2次ピークが発生するが、標準回帰のように$psim n$ではなく$psim nm$となる。
この過度な適合は、トレーニングの暗黙の規則化が完全に作業中であるにもかかわらず、真のスコアではなく、経験的なスコアに向かってモデルを駆動するからである。
- 参考スコア(独自算出の注目度): 5.939780039158003
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Conventional wisdom in deep learning holds that overparameterization---having more parameters $p$ than training samples $n$---is benign: larger models generalize better and, even without regularization, interpolating models generalize well, the test error following a double-descent curve. One might expect the same benign overfitting for diffusion models, whose training reduces to regression, i.e. to minimizing a quadratic score-matching loss. Yet the opposite is observed: overfitting here is catastrophic, driving the model into a memorization regime. We resolve this paradox by combining experiments on U-Nets trained on CelebA with a random-features model for which we derive closed-form learning curves. We show that with a fixed number $m$ of noise realizations per training sample, an interpolation peak does occur, but at $p\sim nm$ rather than at $p\sim n$ as in standard regression. The rise of the test loss, however, sets in much earlier, at $p\sim n$, independently of $m$. This overfitting is malign because, although the implicit regularization of training is fully at work, it drives the model toward the empirical score, which memorizes the training set, rather than toward the true score. A bias-variance decomposition pinpoints the mechanism: the bias of the score estimator starts to grow at $p\sim n$; past the peak the variance decays, as in regression, whereas the bias keeps growing and both saturate at a large value. Since diffusion models are trained with $m\gg1$, the peak is pushed to very large model sizes, and therefore sit on the rising branch that precedes it, where malign overfitting is already in play. Nevertheless, overparameterization remains beneficial when paired with regularization: in the random-features theory and in U-Net experiments, optimally regularized large models---via a ridge penalty or early stopping, respectively---outperform any unregularized models.
- Abstract(参考訳): 従来のディープラーニングの知恵は、オーバーパラメータ化 -- トレーニングサンプルよりも多くのパラメータを持つ $n$--- は良し悪し: 大きなモデルはより良く一般化し、正規化なしでも、補間モデルはうまく一般化する。
二次的なスコアマッチング損失を最小限に抑えるために、トレーニングが回帰に還元される拡散モデルに対して、同じ良質なオーバーフィッティングが期待できるかもしれない。
しかし、反対に、ここでの過度な適合は破滅的であり、モデルを記憶体制へと駆り立てる。
このパラドックスは、CelebAで訓練されたU-Netと、閉形式学習曲線を導出するランダム関数モデルを組み合わせることで解決する。
トレーニングサンプルあたりの雑音実現の固定数$m$では補間ピークが生じるが、標準回帰のように$p\sim n$ではなく$p\sim nm$となる。
しかし、テスト損失の上昇は、$m$とは独立に$p\sim n$をはるかに早く設定する。
この過度な適合は、トレーニングの暗黙の規則化が完全に作業中であるにもかかわらず、モデルが真のスコアではなく、トレーニングセットを記憶する経験的なスコアに向かっているためである。
スコア推定器のバイアスは$p\sim n$で成長し始め、ピークを過ぎると、回帰のように分散は減衰するが、バイアスは成長し続け、どちらも大きな値で飽和する。
拡散モデルは$m\gg1$で訓練されるので、ピークは非常に大きなモデルサイズにプッシュされるため、その前に上昇する分岐の上に座る。
それでも、過パラメータ化は正規化と組み合わせて有用であり、ランダム関数理論やU-Net実験では、尾根のペナルティや早期停止によって、任意の非正規化モデルを上回る最適に正規化された大モデルである。
関連論文リスト
- Learning Shrinks the Hard Tail: Training-Dependent Inference Scaling in a Solvable Linear Model [2.7074235008521246]
ニューラルネットワークのスケーリング法則を最終層微細チューニングの解法モデルで解析する。
学習がエラー分布の「ハードテール」を小さくすることを示す。
論文 参考訳(メタデータ) (2026-01-07T10:00:17Z) - Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling [34.69440744042684]
推論時間スケーリングを解析的に抽出可能なモデルを導入する。
我々は,これらの事実を大言語モデル推論で実験的に検証し,さらに大きな言語モデルを判断する。
論文 参考訳(メタデータ) (2025-12-22T22:13:06Z) - Scaling Laws in Linear Regression: Compute, Parameters, and Data [86.48154162485712]
無限次元線形回帰セットアップにおけるスケーリング法則の理論について検討する。
テストエラーの再現可能な部分は$Theta(-(a-1) + N-(a-1)/a)$であることを示す。
我々の理論は経験的ニューラルスケーリング法則と一致し、数値シミュレーションによって検証される。
論文 参考訳(メタデータ) (2024-06-12T17:53:29Z) - Rejection via Learning Density Ratios [50.91522897152437]
拒絶による分類は、モデルを予測しないことを許容する学習パラダイムとして現れます。
そこで我々は,事前学習したモデルの性能を最大化する理想的なデータ分布を求める。
私たちのフレームワークは、クリーンでノイズの多いデータセットで実証的にテストされます。
論文 参考訳(メタデータ) (2024-05-29T01:32:17Z) - Towards an Understanding of Benign Overfitting in Neural Networks [104.2956323934544]
現代の機械学習モデルは、しばしば膨大な数のパラメータを使用し、通常、トレーニング損失がゼロになるように最適化されている。
ニューラルネットワークの2層構成において、これらの良質な過適合現象がどのように起こるかを検討する。
本稿では,2層型ReLUネットワーク補間器を極小最適学習率で実現可能であることを示す。
論文 参考訳(メタデータ) (2021-06-06T19:08:53Z) - Optimization Variance: Exploring Generalization Properties of DNNs [83.78477167211315]
ディープニューラルネットワーク(DNN)のテストエラーは、しばしば二重降下を示す。
そこで本研究では,モデル更新の多様性を測定するために,新しい測度である最適化分散(OV)を提案する。
論文 参考訳(メタデータ) (2021-06-03T09:34:17Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。