論文の概要: Optimal Stratified Allocation for Rare-Event Onset Forecasting in Dependent Sequences
- arxiv url: http://arxiv.org/abs/2609.04420v1
- Date: Thu, 03 Sep 2026 19:38:01 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-07 18:15:23.804911
- Title: Optimal Stratified Allocation for Rare-Event Onset Forecasting in Dependent Sequences
- Title(参考訳): 従属系列における希少事象発生予測のための最適階層配置法
- Abstract要約: 我々は,K0とK1を2つの層に配置した設計の下で,サブサンプルKnから全体のリスクを推定する。
重み付きリスク推定器の厳密な有限個体群分散をクラス条件によるサンプリングで導出する。
- 参考スコア(独自算出の注目度): 2.2632368327435732
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Let a finite population of n labelled examples carry a class-weighted loss, with pi*n in a rare positive class weighted by N0/N1. We study estimation of total risk from a subsample K << n under designs allocating K0 and K1 draws to the two strata. We derive the exact finite-population variance of the weighted risk estimator under class-conditional sampling without replacement and solve for the optimal allocation. The class multiplier inflates positive-stratum dispersion by the imbalance ratio, causing that ratio to cancel from the optimal allocation and making equal, rather than proportional, allocation the natural default. Simple random sampling is dominated by an explicit between-stratum term; an exact bias identity shows that cluster-representative selection has no general unbiasedness guarantee; and a Serfling bound transfers the allocation result to selection error over a finite candidate set. Under the implemented truncation, the realised allocation ratio is gamma=min(2*pi/f,1), independent of n, yielding the parameter-free efficiency prediction A(pi,f)=gamma/[pi(1-pi)(1+gamma)^2]. A separate measurability result bounds the record occupied by a labelled example, determining the required train-test separation and controlling departure from block independence under absolute regularity. We test these predictions on forecasting the onset of statistically explosive price regimes, dated ex post by the Phillips-Shi-Yu procedure, using 350 U.S. equities from 2004-2011 with under 1% positive rows and five purged forward blocks. The predicted ordering of the four designs holds, and at the 10-day horizon the five design points are ordered exactly as predicted by A (Spearman rho=1, exact p=0.0167). The predicted dependence on pi across horizons does not hold; we identify the channels lying outside the design-based argument.
- Abstract(参考訳): 有限個の n 個のラベル付き例を、N0/N1 で重み付けられた稀な正のクラスにおける pi*n のクラス重み付き損失とする。
K0 と K1 を 2 つの層に配置した設計の下で, サブサンプル K<<n> から総リスクを推定する。
我々は,クラス条件付きサンプリングにおける重み付きリスク推定器の正確な有限集団分散を導出し,最適な割り当てを解く。
クラス乗算器は不均衡比で正層分散を膨らませ、その比は最適な割り当てからキャンセルされ、比例ではなく自然のデフォルトに等しい。
単純なランダムサンプリングは、明示的な層間項によって支配される; 正確なバイアスIDは、クラスタ表現の選択が一般的な不偏性保証を持たないことを示し、Serfling境界は、有限候補集合上の選択誤差に割り当て結果を転送する。
実現された割当比は、n に依存しないγ=min(2*pi/f,1) であり、パラメータフリーな効率予測 A(pi,f)=gamma/[pi(1-pi)(1+gamma)^2] が得られる。
別個の測定可能性の結果は、ラベル付き例によって占有された記録を束縛し、必要な列車-テスト分離を決定し、絶対正則性の下でブロック独立からの離脱を制御する。
我々は,2004~2011年の米国株式350株を1%以下の正の行と5つの前向きブロックを用いて,Philips-Shi-Yu法により記載された,統計的に爆発的な価格体系の開始を予測した上で,これらの予測を検証した。
4つのデザインの予測順序は成り立ち、10日間の水平線では、5つのデザインポイントはA(スピアマンrho=1, 正確なp=0.0167)によって予測されるように順序付けられている。
水平線を越えたπの予測依存性は保たず、設計に基づく議論の外にあるチャネルを識別する。
関連論文リスト
- Semiparametric Inference for Conditional Shapley Feature Importance [0.0]
本稿では, 条件定式化について検討し, その実条件分布の下で, 対数外特徴を積分する。
そこで本研究では, K-fold cross-fitting とU-statistic correct of the squared loss intervalを提案する。
ピンスカー型境界は、作業コプラクラスを誤って特定するバイアスを定量化し、ワインコプラは条件付きサンプリングが可能である。
論文 参考訳(メタデータ) (2026-09-09T15:19:24Z) - What Fixed-Rollout pass@k Evaluations Can Identify [3.0318216701273397]
繰り返しサンプリング評価は、問題毎に収集されたサンプル数nをはるかに超え、pass@kを外挿する。
プール/ランダム・タスク条件付きビノミカルモデルでは、固定nの成功回数は、潜在タスク毎の成功分布の n 自由モーメントのみを識別する。
論文 参考訳(メタデータ) (2026-09-08T10:27:52Z) - One-step lowest-variance selection in a Gaussian random-field model motivated by masked diffusion: Total correlation and a square root collision threshold [8.230831152671708]
マスク付き離散拡散における信頼誘導型並列アンマスキングの動機付けにより、スタイリングされたガウス確率場モデルにおける1つの選択ステップについて検討する。
その結果、マスク付き離散拡散における厳密な信頼幾何と空間相関が一段階選択に結びついている。
論文 参考訳(メタデータ) (2026-07-20T03:56:34Z) - K-ABENA: K-Adaptive Backpropagation with Error-based N-exclusion Algorithm : (Compensated Loss-Based Sample Exclusion with Unbiased Gradient Estimation) [0.0]
K-ABENAK-Adaptive Backpropagation with Error-based N-exclusionは選択的な計算フレームワークである。
これは、少量の低損失("ミニ")観測を除外することで、イット当たりの計算トレーニングコストを削減する。
ラベル制限下で0.386の精度(ベースライン:0.832)、極端な不均衡下で0.53のAUCに崩壊する。
論文 参考訳(メタデータ) (2026-07-07T06:58:51Z) - How Useful is Causal Invariance for Domain Adaptation in Finite-Sample Settings? [58.740078141879984]
機械学習モデルは、トレーニングされたソースディストリビューションとは異なるターゲットディストリビューションにデプロイされると、しばしば劣化する。
因果関係に基づく領域一般化における最近の研究は、共用因果構造が不変な予測因子を誘導する方法を示している。
本稿では,完全あるいは部分的な因果知識が,教師付きドメイン適応を確実に改善できるかどうかについて検討する。
論文 参考訳(メタデータ) (2026-06-10T21:07:49Z) - Beyond Uncertainty Sets: Leveraging Optimal Transport to Extend Conformal Predictive Distribution to Multivariate Settings [0.14504054468850666]
コンフォーマル予測(CP)は、有限サンプルカバレッジを保証するモデル出力に対する不確実性集合を構成する。
最適な割り当ては、スコア空間の固定された多面体分割において一括的に一定であることを示す。
これにより、予測セット全体を魅力的に特徴付けることができ、予測セットのより深い制限に対処するための機械を提供する。
論文 参考訳(メタデータ) (2025-11-19T05:59:01Z) - Theory on Score-Mismatched Diffusion Models and Zero-Shot Conditional Samplers [49.97755400231656]
一般のスコアミスマッチ拡散サンプリング器に対する明示的な次元依存性を持つ最初の性能保証を示す。
その結果, スコアミスマッチは, 目標分布とサンプリング分布の分布バイアスとなり, 目標分布とトレーニング分布の累積ミスマッチに比例することがわかった。
この結果は、測定ノイズに関係なく、任意の条件モデルに対するゼロショット条件付きサンプリングに直接適用することができる。
論文 参考訳(メタデータ) (2024-10-17T16:42:12Z) - Relaxed Quantile Regression: Prediction Intervals for Asymmetric Noise [51.87307904567702]
量子レグレッション(Quantile regression)は、出力の分布における量子の実験的推定を通じてそのような間隔を得るための主要なアプローチである。
本稿では、この任意の制約を除去する量子回帰に基づく区間構成の直接的な代替として、Relaxed Quantile Regression (RQR)を提案する。
これにより、柔軟性が向上し、望ましい品質が向上することが実証された。
論文 参考訳(メタデータ) (2024-06-05T13:36:38Z) - Conformal Language Modeling [61.94417935386489]
生成言語モデル(LM)の共形予測のための新しい手法を提案する。
標準共形予測は厳密で統計的に保証された予測セットを生成する。
我々は,オープンドメイン質問応答,テキスト要約,ラジオロジーレポート生成において,複数のタスクに対するアプローチの約束を実証する。
論文 参考訳(メタデータ) (2023-06-16T21:55:08Z) - Distribution-Free Robust Linear Regression [5.532477732693]
共変体の分布を仮定せずにランダムな設計線形回帰を研究する。
最適部分指数尾を持つオーダー$d/n$の過大なリスクを達成する非線形推定器を構築する。
我々は、Gy"orfi, Kohler, Krzyzak, Walk が原因で、truncated least squares 推定器の古典的境界の最適版を証明した。
論文 参考訳(メタデータ) (2021-02-25T15:10:41Z) - Suboptimality of Constrained Least Squares and Improvements via
Non-Linear Predictors [3.5788754401889014]
有界ユークリッド球における正方形損失に対する予測問題と最良の線形予測器について検討する。
最小二乗推定器に対する$O(d/n)$過剰リスク率を保証するのに十分な分布仮定について論じる。
論文 参考訳(メタデータ) (2020-09-19T21:39:46Z) - Distributional Reinforcement Learning via Moment Matching [54.16108052278444]
ニューラルネットワークを用いて各戻り分布から統計量の有限集合を学習する手法を定式化する。
我々の手法は、戻り分布とベルマン目標の間のモーメントの全ての順序を暗黙的に一致させるものとして解釈できる。
Atariゲームスイートの実験により,本手法は標準分布RLベースラインよりも優れていることが示された。
論文 参考訳(メタデータ) (2020-07-24T05:18:17Z) - Distributionally Robust Bayesian Quadrature Optimization [60.383252534861136]
確率分布が未知な分布の不確実性の下でBQOについて検討する。
標準的なBQOアプローチは、固定されたサンプル集合が与えられたときの真の期待目標のモンテカルロ推定を最大化する。
この目的のために,新しい後方サンプリングに基づくアルゴリズム,すなわち分布的に堅牢なBQO(DRBQO)を提案する。
論文 参考訳(メタデータ) (2020-01-19T12:00:33Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。