論文の概要: Learning from Uncertainty-dependent Missing Labels for Semi-supervised Classification
- arxiv url: http://arxiv.org/abs/2608.23960v1
- Date: Tue, 25 Aug 2026 01:40:47 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-26 14:09:34.728705
- Title: Learning from Uncertainty-dependent Missing Labels for Semi-supervised Classification
- Title(参考訳): 半教師付き分類のための不確実性依存型欠落ラベルからの学習
- Authors: You-Gan Wang, Jinran Wu, Geoffrey J. McLachlan,
- Abstract要約: ラベルの欠落確率が観測された特徴に依存する半教師付き環境について検討する。
このような不確実性に依存しないラベルに対する可能性に基づく情報理論を開発した。
- 参考スコア(独自算出の注目度): 2.5794915063815664
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Missing labels are usually regarded as a source of information loss in classification. We study a semi-supervised setting in which the probability of label missingness depends on the observed features through posterior classification uncertainty. In this setting, the missingness indicator is not only a record of an unobserved label, but also an observable signal generated by a mechanism linked to the classifier. We develop a likelihood-based information theory for such uncertainty-dependent missing labels. Under correct specification, we derive a Fisher-information decomposition that separates a partial-labeling component from a nonnegative mechanism-curvature term. Under joint misspecification of the label model and the missingness mechanism, we obtain the corresponding Godambe--Eicker--Huber--White sensitivity and sandwich-covariance partitions. We also clarify the relevant complete-data benchmark: favorable missingness can increase information relative to ordinary fully labeled or budget-matched non-informative labeling baselines, but cannot exceed the information in the augmented experiment in which labels and mechanism indicators are both observed. For plug-in classifiers, we connect the information decomposition to margin-based excess-risk bounds. In regular two-component mixture settings this yields the parametric \(n^{-1}\) excess-risk rate, with constants determined by the nuisance-adjusted information in discriminant directions. Gaussian-mixture calculations and a medical diagnosis example illustrate how uncertainty-dependent labeling mechanisms can improve estimation and classification under a fixed labeling budget.
- Abstract(参考訳): 失われたラベルは通常、分類における情報損失の源と見なされる。
ラベルの欠落の確率は後続の分類の不確実性によって観測された特徴に依存する半教師付き環境について検討した。
この設定では、欠落指示器は、観測されていないラベルの記録だけでなく、分類器にリンクされたメカニズムによって生成された観測可能な信号でもある。
このような不確実性に依存しないラベルに対する可能性に基づく情報理論を開発した。
正確な仕様の下では、部分ラベル成分を非負の機構曲率項から分離するフィッシャー情報分解を導出する。
ラベルモデルと不足機構の相違により,対応するGodambe--Eicker--Huber--White感度とサンドイッチ共分散分割が得られる。
また, 有意な欠落は, 通常の完全ラベル付きまたは予算に適合しないラベル付けベースラインに対して情報を増やすことができるが, ラベルとメカニズムインジケータの両方が観察される付加実験では, 情報を超えることはできない。
プラグイン分類器に対しては、情報分解をマージンベースの過剰リスク境界に接続する。
通常の2成分混合設定では、これはパラメトリックな \(n^{-1}\) 過剰リスク率をもたらす。
ガウス混合計算と医学診断例は、不確実性に依存したラベル付け機構が、固定されたラベル付け予算の下でどのように評価と分類を改善するかを示す。
関連論文リスト
- Probabilistic Label Spreading: Efficient and Consistent Estimation of Soft Labels with Epistemic Uncertainty on Graphs [4.480864309234644]
本稿では,ラベルの不確かさを推定する確率的ラベル拡散手法を提案する。
データポイントあたりのアノテーションの数が0に収束しても、ラベルの拡散が一貫した確率推定器が得られることを示す。
実験結果から,本手法はベースラインと比較して,所望のラベル品質を実現するために必要なアノテーション予算を大幅に削減することが示された。
論文 参考訳(メタデータ) (2026-02-04T14:00:30Z) - Informative missingness and its implications in semi-supervised learning [2.5794915063815664]
半教師付き学習(SSL)はラベル付きデータと非ラベル付きデータの両方を用いて分類器を構成する。
これは、有限混合モデルに対する可能性フレームワーク内で統計的に定式化できる不完全データ問題を定義する。
このような情報不足をモデル化することは、実証的なSSLメソッドの振る舞いと可能性に基づく推論を統一するコヒーレントな統計フレームワークを提供する。
論文 参考訳(メタデータ) (2025-12-04T02:26:56Z) - SSLfmm: An R Package for Semi-Supervised Learning with a Mixed-Missingness Mechanism in Finite Mixture Models [2.0253523660913664]
半教師付き学習(SSL)は、観測のサブセットのみをラベル付けしたデータセットから分類器を構築する。
観察が損なわれない可能性は、その特徴ベクトルのあいまいさに依存する可能性があるため、不足過程は有益なものとなる。
このパッケージにはモデリングの実用的なツールが含まれており、シミュレートされた例を通してそのパフォーマンスを説明している。
論文 参考訳(メタデータ) (2025-12-03T00:14:33Z) - Label Distribution Learning with Biased Annotations by Learning Multi-Label Representation [120.97262070068224]
マルチラベル学習(MLL)は,実世界のデータ表現能力に注目されている。
ラベル分布学習(LDL)は正確なラベル分布の収集において課題に直面している。
論文 参考訳(メタデータ) (2025-02-03T09:04:03Z) - Dist-PU: Positive-Unlabeled Learning from a Label Distribution
Perspective [89.5370481649529]
本稿では,PU学習のためのラベル分布視点を提案する。
そこで本研究では,予測型と基底型のラベル分布間のラベル分布の整合性を追求する。
提案手法の有効性を3つのベンチマークデータセットで検証した。
論文 参考訳(メタデータ) (2022-12-06T07:38:29Z) - How Does Pseudo-Labeling Affect the Generalization Error of the
Semi-Supervised Gibbs Algorithm? [73.80001705134147]
擬似ラベル付き半教師付き学習(SSL)におけるGibsアルゴリズムによる予測一般化誤差(ゲンエラー)を正確に評価する。
ゲンエラーは、出力仮説、擬ラベルデータセット、ラベル付きデータセットの間の対称性付きKL情報によって表現される。
論文 参考訳(メタデータ) (2022-10-15T04:11:56Z) - Incorporating Label Uncertainty in Understanding Adversarial Robustness [17.65850501514483]
最先端モデルによって誘導される誤差領域は、ランダムに選択されたサブセットよりもラベルの不確実性が高い傾向を示す。
この観測は,ラベルの不確実性を考慮した濃度推定アルゴリズムの適用を動機付けている。
論文 参考訳(メタデータ) (2021-07-07T14:26:57Z) - Distribution-free uncertainty quantification for classification under
label shift [105.27463615756733]
2つの経路による分類問題に対する不確実性定量化(UQ)に焦点を当てる。
まず、ラベルシフトはカバレッジとキャリブレーションの低下を示すことでuqを損なうと論じる。
これらの手法を, 理論上, 分散性のない枠組みで検討し, その優れた実用性を示す。
論文 参考訳(メタデータ) (2021-03-04T20:51:03Z) - Comparing the Value of Labeled and Unlabeled Data in Method-of-Moments
Latent Variable Estimation [17.212805760360954]
我々は,メソッド・オブ・モーメント・潜在変数推定におけるモデル誤特定に着目したフレームワークを用いている。
そして、ある場合においてこのバイアスを確実に排除する補正を導入する。
理論上, 合成実験により, 特定されたモデルではラベル付点がラベル付点以上の定数に値することを示した。
論文 参考訳(メタデータ) (2021-03-03T23:52:38Z) - Exploiting Sample Uncertainty for Domain Adaptive Person
Re-Identification [137.9939571408506]
各サンプルに割り当てられた擬似ラベルの信頼性を推定・活用し,ノイズラベルの影響を緩和する。
不確実性に基づく最適化は大幅な改善をもたらし、ベンチマークデータセットにおける最先端のパフォーマンスを達成します。
論文 参考訳(メタデータ) (2020-12-16T04:09:04Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。