論文の概要: Two-Sample Testing via Path-based Inference
- arxiv url: http://arxiv.org/abs/2610.05684v1
- Date: Mon, 05 Oct 2026 01:56:28 GMT
- ステータス: 情報取得中
- システム内更新日: 2026-10-06 20:26:09.441237
- Title: Two-Sample Testing via Path-based Inference
- Title(参考訳): 経路ベース推論による2サンプルテスト
- Abstract要約: 2サンプルテストは、同じ分布が2つの有限データセットを生成するかどうかを決定する問題である。
補間子を用いて、両分布を共有ガウスボトルネックに接続する。
ゼロ仮説は、集団が、または同等に、2つのハーフの速度場が、任意のノイズレベルに一致する場合にのみ成立することを示す。
- 参考スコア(独自算出の注目度): 54.590623065123076
- License:
- Abstract: Modern deep generative models are primarily studied for their ability to generate realistic samples, yet the generative dynamics they learn can also serve as objects of statistical inference. We develop this idea for two-sample testing, the problem of deciding whether the same distribution generated two finite datasets. Using stochastic interpolants, we connect both distributions to a shared Gaussian bottleneck, so that each half of the resulting path is a Gaussian channel acting on a single population. We prove that the null hypothesis holds if and only if the population denoiser, or equivalently, the velocity fields of the two halves, coincide at any single noise level, which amounts to a reflection symmetry of the path about the bottleneck. Deviations from this symmetry yield a continuum of two-sample witnesses, which we estimate via held-out regression risks on learned denoisers and velocities and aggregate along the path; under an information-theoretic weighting, the aggregated discrepancy equals the Jeffreys divergence between the noise-smoothed distributions. Calibrating the resulting statistics by permutation yields tests that are valid in finite samples for any trained networks and consistent when the fields are learned accurately. On a synthetic benchmark and three image benchmarks, the proposed tests improve power over the strongest baseline by up to 33 percentage points at an equal total sample budget, with the best choice of regression representation and path weighting depending on the data modality. These results show that generative paths provide a principled representation for statistical testing, extending stochastic-interpolant models beyond generation.
- Abstract(参考訳): 現代の深層生成モデルは、主に現実的なサンプルを生成する能力のために研究されているが、彼らが学習した生成力学は統計的推論のオブジェクトとしても機能する。
本研究では,同じ分布が2つの有限データセットを生成するかどうかを判定する2サンプルテストの考え方を考案した。
確率的補間子を用いて、両分布を共有ガウス的ボトルネックに結び付け、その結果の経路の半数が1つの集団に作用するガウス的チャネルとなる。
ゼロ仮説が、集団が、または同等に、2つのハーフの速度場が、ボトルネックに関する経路の反射対称性に比例する任意の単一ノイズレベルに一致することを証明している。
この対称性から逸脱した2サンプルの証人の連続性は、学習した騒音と速度の保留回帰リスクを通じて推定され、経路に沿って集約される。
置換によって得られた統計を校正すると、任意の訓練されたネットワークの有限サンプルで有効であり、フィールドが正確に学習されたときに一貫したテストが得られる。
合成ベンチマークと3つの画像ベンチマークでは,データモダリティに応じて回帰表現と経路重み付けを最適に選択することにより,最大33ポイントの最大ベースラインに対するパワー向上が図られた。
これらの結果は、生成経路が統計的テストの原理的な表現を提供し、確率補間モデルを生成を超えて拡張することを示している。
関連論文リスト
- Sample complexity bounds for categorical Markov random fields via Discrete Diffusions [2.0339465062586566]
我々は、離散拡散のためのエンドツーエンドのサンプル複雑性境界を持つ学習方法を開発した。
私たちの主要な技術的洞察は、離散的なスコアを分解する新しいエンピンジングです。
両立型ニューラルスコア学習器を提案し,それを$$-leapingと組み合わせて,エンドツーエンドのサンプリング手順を得る。
論文 参考訳(メタデータ) (2026-10-01T17:40:21Z) - Counterfactual Generation via Flow Matching: Coupling-Sensitive End-to-End Rates [11.51425194152931]
本研究では,サンプル分割された2つの頑健なトレーニング目標と,観測結果と適合条件付き結果モデルから抽出した目標結果との学習的結合を組み合わせたフローパラメトリック手法を開発した。
我々の理論的な主な貢献は、定数ステップの離散化のための結合感受性KL結合である。
合成および半合成画像ベンチマークの実験は、結合依存理論をサポートし、有限の離散化予算において、サンプルが対応するODEサンプリング器より優れていることを示す。
論文 参考訳(メタデータ) (2026-10-01T07:05:31Z) - On the Wasserstein Convergence and Straightness of Rectified Flow [54.580605276017096]
Rectified Flow (RF) は、ノイズからデータへの直流軌跡の学習を目的とした生成モデルである。
RFのサンプリング分布とターゲット分布とのワッサーシュタイン距離に関する理論的解析を行った。
本稿では,従来の経験的知見と一致した1-RFの特異性と直線性を保証する一般的な条件について述べる。
論文 参考訳(メタデータ) (2024-10-19T02:36:11Z) - Theory on Score-Mismatched Diffusion Models and Zero-Shot Conditional Samplers [49.97755400231656]
一般のスコアミスマッチ拡散サンプリング器に対する明示的な次元依存性を持つ最初の性能保証を示す。
その結果, スコアミスマッチは, 目標分布とサンプリング分布の分布バイアスとなり, 目標分布とトレーニング分布の累積ミスマッチに比例することがわかった。
この結果は、測定ノイズに関係なく、任意の条件モデルに対するゼロショット条件付きサンプリングに直接適用することができる。
論文 参考訳(メタデータ) (2024-10-17T16:42:12Z) - A Kernel-Based Conditional Two-Sample Test Using Nearest Neighbors (with Applications to Calibration, Regression Curves, and Simulation-Based Inference) [3.622435665395788]
本稿では,2つの条件分布の違いを検出するカーネルベースの尺度を提案する。
2つの条件分布が同じである場合、推定はガウス極限を持ち、その分散はデータから容易に推定できる単純な形式を持つ。
また、条件付き適合性問題に適用可能な推定値を用いた再サンプリングベースのテストも提供する。
論文 参考訳(メタデータ) (2024-07-23T15:04:38Z) - Unsupervised Learning of Sampling Distributions for Particle Filters [80.6716888175925]
観測結果からサンプリング分布を学習する4つの方法を提案する。
実験により、学習されたサンプリング分布は、設計された最小縮退サンプリング分布よりも優れた性能を示すことが示された。
論文 参考訳(メタデータ) (2023-02-02T15:50:21Z) - Comparing two samples through stochastic dominance: a graphical approach [2.867517731896504]
実世界のシナリオでは非決定論的測定が一般的である。
推定累積分布関数に従って2つのサンプルを視覚的に比較するフレームワークを提案する。
論文 参考訳(メタデータ) (2022-03-15T13:37:03Z) - Unrolling Particles: Unsupervised Learning of Sampling Distributions [102.72972137287728]
粒子フィルタリングは複素系の優れた非線形推定を計算するために用いられる。
粒子フィルタは様々なシナリオにおいて良好な推定値が得られることを示す。
論文 参考訳(メタデータ) (2021-10-06T16:58:34Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。