論文の概要: Measuring consistency via ensemble margin and local prediction variability: Auditing decision systems in the presence of predictive multiplicity
- arxiv url: http://arxiv.org/abs/2609.01397v1
- Date: Tue, 01 Sep 2026 15:22:01 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-02 16:31:36.79751
- Title: Measuring consistency via ensemble margin and local prediction variability: Auditing decision systems in the presence of predictive multiplicity
- Title(参考訳): アンサンブルマージンと局所予測変数による一貫性の測定:予測多重性の有無による意思決定システムの検討
- Authors: Sinjini Banerjee, Tim Marrinan, Anand D. Sarwate,
- Abstract要約: ラショウモン効果は、同じ精度のモデルが同じ入力に対して異なる予測を生成する機械学習現象である。
本稿では,アンサンブルマージンと各構成モデルに対する局所的予測変数の尺度を組み合わせた一貫性基準を提案する。
実験の結果,Rashomon集合からのアンサンブルモデルにより,誤検出のリスクを大幅に低減できることがわかった。
- 参考スコア(独自算出の注目度): 4.924900955306918
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: The Rashomon effect is a machine learning phenomenon where equally accurate models produce different predictions for the same inputs (predictive multiplicity). Existing work primarily focuses on multiplicity within individual models, but in more complex decision systems, the impact of the Rashomon effect is less well understood. In this work, we study multiplicity from the perspective of auditing incorrect ensemble predictions, where the decision to divert an instance for human review is based on a consistency criterion that combines the ensemble margin with a measure of local prediction variability for each constituent model. With mild assumptions about stability and smoothness, we show that the consistency scores of finite ensembles converge to the corresponding consistency score of the expected model from the Rashomon set as the ensemble size and the number of samples used to measure local prediction variability increase. To demonstrate the efficacy of the proposed criterion, we evaluate the framework with respect to transformer models applied to natural language understanding tasks and parameter-efficient fine-tuning of large language models used for tabular data classification tasks. Our experiments show that ensembling models from the Rashomon set substantially reduces the risk of incorrect predictions going unchecked compared with auditing a single model, while incurring only a moderate increase in the number of diversions. Moreover, the auditing behavior of the full Rashomon set can be closely approximated by finite ensembles of relatively modest size, with the risk approaching zero for some datasets. We further demonstrate that the proposed measure exhibits stronger agreement with established predictive multiplicity metrics than existing consistency measures, providing a more reliable way to capture multiplicity in the Rashomon set.
- Abstract(参考訳): ラショモン効果(英: Rashomon effect)は、同じモデルが同じ入力(予測多重性)に対して異なる予測を生成する機械学習現象である。
既存の研究は主に個々のモデル内の多重性に焦点を当てているが、より複雑な決定システムでは、ラショモン効果の影響は理解されていない。
本研究では,不正確なアンサンブル予測を監査する視点から,アンサンブルマージンと各構成モデルの局所的予測変数の尺度を組み合わせた一貫性基準に基づいて,人間のレビューのためにインスタンスを分割する決定を行う。
安定性と滑らか性に関する軽度な仮定により、有限アンサンブルの一貫性スコアは、ラッショモン集合から期待されるモデルの一貫性スコアにアンサンブルサイズとして収束し、局所的な予測変数の増大を測定するために使用されるサンプルの数が増加することを示す。
提案手法の有効性を示すために, 自然言語理解タスクに適用されたトランスフォーマーモデルと, 表型データ分類タスクに使用される大規模言語モデルのパラメータ効率の高い微調整について, フレームワークの評価を行った。
実験の結果, ラッショモン集合からのアンサンブルモデルでは, 単一モデルの監査に比べて誤判定のリスクが大幅に低減され, 転化回数の適度な増加がみられた。
さらに、完全な羅生門集合の監査挙動は、比較的控えめな大きさの有限アンサンブルによって近似され、いくつかのデータセットに対してゼロに近づく危険性がある。
さらに,提案手法は,既存の整合性尺度よりも確立された予測多重度指標との強い一致を示し,ラショモン集合の多重度を捉えるための信頼性の高い手段を提供する。
関連論文リスト
- Mitigating the Multiplicity Burden: The Role of Calibration in Reducing Predictive Multiplicity of Classifiers [0.0]
本稿では,分類校正と予測乗算の相互作用について検討する。
マイノリティクラスの観察は、不均等な多種多様性の重荷を負う。
ポストホックキャリブレーション法の適用は、ラショモン集合全体の低視認性と関連している。
論文 参考訳(メタデータ) (2026-03-12T09:54:07Z) - On Arbitrary Predictions from Equally Valid Models [49.56463611078044]
モデル多重性(英: Model multiplicity)とは、同じ患者に対して矛盾する予測を認める複数の機械学習モデルを指す。
たとえ小さなアンサンブルであっても、実際は予測的多重性を緩和・緩和できることを示す。
論文 参考訳(メタデータ) (2025-07-25T16:15:59Z) - Regularized Neural Ensemblers [55.15643209328513]
本研究では,正規化ニューラルネットワークをアンサンブル手法として活用することを検討する。
低多様性のアンサンブルを学習するリスクを動機として,ランダムにベースモデル予測をドロップすることで,アンサンブルモデルの正規化を提案する。
このアプローチはアンサンブル内の多様性の低い境界を提供し、過度な適合を減らし、一般化能力を向上させる。
論文 参考訳(メタデータ) (2024-10-06T15:25:39Z) - Quantifying Prediction Consistency Under Fine-Tuning Multiplicity in Tabular LLMs [10.494477811252034]
微調整多重度は分類タスクにおけるタブラル LLM に現れる。
我々の研究は、タブラルLLMにおける微調整多重性というこのユニークな挑戦を定式化する。
本稿では,コストのかかるモデル再訓練を伴わずに,個々の予測の一貫性を定量化する手法を提案する。
論文 参考訳(メタデータ) (2024-07-04T22:22:09Z) - Multi-View Conformal Learning for Heterogeneous Sensor Fusion [0.12086712057375555]
異種センサ融合のためのマルチビュー・シングルビューコンフォメーションモデルの構築と試験を行った。
我々のモデルは、共形予測フレームワークに基づいているため、理論的な限界信頼保証を提供する。
また,複数ビューモデルが単一ビューモデルに比べて不確実性の低い予測セットを生成することを示した。
論文 参考訳(メタデータ) (2024-02-19T17:30:09Z) - Predictive Churn with the Set of Good Models [61.00058053669447]
本稿では,予測的不整合という2つの無関係な概念の関連性について考察する。
予測多重性(英: predictive multiplicity)は、個々のサンプルに対して矛盾する予測を生成するモデルである。
2つ目の概念である予測チャーン(英: predictive churn)は、モデル更新前後の個々の予測の違いを調べるものである。
論文 参考訳(メタデータ) (2024-02-12T16:15:25Z) - Dropout-Based Rashomon Set Exploration for Efficient Predictive
Multiplicity Estimation [15.556756363296543]
予測多重性(英: Predictive multiplicity)とは、ほぼ等しい最適性能を達成する複数の競合モデルを含む分類タスクを指す。
本稿では,Rashomon 集合のモデル探索にドロップアウト手法を利用する新しいフレームワークを提案する。
本手法は, 予測多重度推定の有効性の観点から, ベースラインを一貫して上回ることを示す。
論文 参考訳(メタデータ) (2024-02-01T16:25:00Z) - Trusted Multi-View Classification [76.73585034192894]
本稿では,信頼された多視点分類と呼ばれる新しい多視点分類手法を提案する。
さまざまなビューをエビデンスレベルで動的に統合することで、マルチビュー学習のための新しいパラダイムを提供する。
提案アルゴリズムは,分類信頼性とロバスト性の両方を促進するために,複数のビューを併用する。
論文 参考訳(メタデータ) (2021-02-03T13:30:26Z) - Characterizing Fairness Over the Set of Good Models Under Selective
Labels [69.64662540443162]
同様の性能を実現するモデルセットに対して,予測公正性を特徴付けるフレームワークを開発する。
到達可能なグループレベルの予測格差の範囲を計算するためのトラクタブルアルゴリズムを提供します。
選択ラベル付きデータの実証的な課題に対処するために、我々のフレームワークを拡張します。
論文 参考訳(メタデータ) (2021-01-02T02:11:37Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。