論文の概要: Reliable Benchmarking of Artifact Detection in Computational Pathology: A Reproducibility and Uncertainty Analysis
- arxiv url: http://arxiv.org/abs/2608.30835v1
- Date: Mon, 31 Aug 2026 14:07:15 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-01 18:31:31.434729
- Title: Reliable Benchmarking of Artifact Detection in Computational Pathology: A Reproducibility and Uncertainty Analysis
- Title(参考訳): 計算病理におけるアーティファクト検出の信頼性ベンチマーク:再現性と不確かさ解析
- Authors: Konstantinos Moutselos, Ilias Maglogiannis,
- Abstract要約: メソッド: このプロトコルは、テストセットサンプリング、トレーニング性、パーティション構成、文書化前処理の4つの変数ソースを定量化する。
拡散型人工物検出装置を独立に再構築し, 元の24スライディング分割と教師付きベースラインに対して評価した。
結果: 本手法の中枢機構は、補助的コントラスト項は、プールされたF1を0.673から0.688に改善し、第2シードで複製する。
- 参考スコア(独自算出の注目度): 1.8864137667201766
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Background and Objective: Quality control is a prerequisite for whole-slide image analysis, yet the benchmarks on which quality-control methods are compared share four properties that make their reported differences hard to interpret: few independent slides, annotation concentrated in a minority of them, pooled ratio metrics with no closed-form standard error, and a single inherited train/test partition. We propose a reliability protocol for such benchmarks. Methods: The protocol quantifies four sources of variability - test-set sampling, training stochasticity, partition composition, and undocumented preprocessing - a claim is reportable only if it survives all four; three of the four cost minutes of compute. We apply it to an independent reconstruction of a published diffusion-based artifact detector, evaluated on the original 24-slide partition and against a supervised baseline. Results: The method's central mechanism reproduces: the auxiliary contrastive term improves pooled F1 from 0.673 to 0.688 and replicates under a second seed (+0.0156, p = 0.031; +0.0190, p = 0.005), although it acts on pen marking rather than the artifact types cited to motivate it. Its comparative claims do not: differences between design variants, and against the supervised baseline, fall inside the uncertainty of the evaluation. Four of 24 slides carry 70% of scored annotated pixels, giving an effective sample size of 6.2, and the inherited partition sits at the 7th percentile. An unreported tissue-restriction step excludes 41.4% of out-of-focus annotation against 2.6% of air bubble; such a gate is confounded with blur by construction. Conclusions: Small-cohort benchmarks support far weaker conclusions than current reporting implies. The four checks are cheap enough to accompany any evaluation on such a resource and separate reproducible effects from differences the evaluation cannot resolve.
- Abstract(参考訳): 背景と目的: 品質管理は、全体スライディング画像分析の前提条件であるが、品質管理手法を比較したベンチマークでは、報告された違いを解釈しにくくする4つの特性を共有している。
このようなベンチマークのための信頼性プロトコルを提案する。
メソッド: このプロトコルは、テストセットサンプリング、トレーニング確率性、パーティション構成、未文書の事前処理の4つの変数ソースを定量化します。
拡散型人工物検出装置を独立に再構築し, 元の24スライディング分割と教師付きベースラインに対して評価した。
結果: 補助的コントラスト項は、プールされたF1を0.673から0.688に改善し、第2シード(+0.0156, p = 0.031; +0.0190, p = 0.005)で複製する。
その比較主張はそうではない: 設計の変種の違いと、教師付きベースラインの違いは、評価の不確実性に陥る。
24枚のスライドのうち4枚は70%のアノテート画素を持ち、有効サンプルサイズは6.2で、継承されたパーティションは7番目のパーセンタイルにある。
報告されていない組織制限ステップでは、アウト・オブ・フォーカスのアノテーションの41.4%をエアバブルの2.6%から除外している。
結論: 小規模コホートベンチマークは、現在の報告が示すよりもはるかに弱い結論を支持する。
4つのチェックは、そのようなリソースに対する評価に付随するほど安価であり、再現可能な効果を、評価が解決できない相違から分離する。
関連論文リスト
- Decomposing Wrong-Consensus Agreement in LLM Self-Consistency: A GPT-4.1 Case Study [1.2716872085463884]
複数のLSMサンプルに対する多数決投票は、回答の正確性を高めるために広く用いられているが、その利得は不規則に異なる。
本稿は、この失敗を定量的に分析する。
論文 参考訳(メタデータ) (2026-08-19T10:50:15Z) - Automated binary classification of hazelnut X-ray images: A deep-learning benchmark for quality assessment [33.72751145910978]
非破壊的なX線撮影では、外部検査だけでは検出が難しい内部のハゼルナッツの欠陥が明らかになる。
799セグメントのシングルカーネルX線画像に基づくバイナリ・ヘーゼルナットの品質分類のベンチマークを提案する。
論文 参考訳(メタデータ) (2026-08-12T07:55:22Z) - Passing Coarse Marginal Checks Can Be Cheap: Persona Mixtures and Imprecise Treatment-Response Estimates in an LLM Persona Panel [0.0]
戦略ゲームにおける16個の軽量人格条件付きGPT-4.1構成の固定パネルについて検討した。
変化は急激なインデクシングであったが、そのシェアは不確実性の仮定に依存していた。
結果は、1つの固定されたモデル・プロンプトパネルに関係し、人間の置換性を確立しない。
論文 参考訳(メタデータ) (2026-08-02T04:10:11Z) - Jury Duty: Calibration and Orientation Failures in MLLM-as-a-Judge Under Cultural Ambiguity [43.23157850664282]
VOIR DIREは、米国と中国本土にまたがる、626の文化的にペアリングされたイメージプロンプトアーティファクトのマルチモーダルベンチマークである。
6つのMLLMにまたがって、バイアスは2つの障害に分解される。
本稿では,各参照プールに対して個別にアライメントを報告することを推奨する。
論文 参考訳(メタデータ) (2026-06-12T15:53:40Z) - Distributional Energy-Based Models for Uncertainty-Aware Structured LLM Reasoning [40.342912574072024]
大規模言語モデルは、旅行計画やコードソリューションのような構造化されたアウトプットを生成する。
個々の推論ステップは正しく見えるが、アウトプット全体が予算に違反したり、テストケースに失敗したり、あるいは以前の推論に矛盾することがある。
構造化LCM出力の検証のための決定論的解析制約付き学習品質スコアラを提案する。
論文 参考訳(メタデータ) (2026-05-15T17:08:27Z) - PAIR-CI: Calibrated Conditional Independence Testing for Causal Discovery with Incomplete Data [0.0]
PAIR-CIは非パラメトリック条件独立(CI)テストであり,複数の命令を直接推論手順に統合することによりキャリブレーションを回復する。
確率的に一貫した分散推定器は、クロスバリデーションと多重計算による不確かさを共同で説明する。
論文 参考訳(メタデータ) (2026-05-06T12:34:37Z) - When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning [16.505918019260964]
信頼性と信頼性の低い予測を混合することにより,最先端モデル(Qwen2.5-Math-7B)の精度が61%向上することが実証された。
正しい予測の18.4%は安定で忠実な推論を採用しており、81.6%は計算的に一貫性のない経路を通して現れる。
論文 参考訳(メタデータ) (2026-03-03T19:43:36Z) - CLUE: Non-parametric Verification from Experience via Hidden-State Clustering [64.50919789875233]
隠れアクティベーションの軌跡内の幾何的に分離可能なシグネチャとして解の正しさが符号化されていることを示す。
ClUE は LLM-as-a-judge ベースラインを一貫して上回り、候補者の再選において近代的な信頼に基づく手法に適合または超えている。
論文 参考訳(メタデータ) (2025-10-02T02:14:33Z) - Causal Understanding by LLMs: The Role of Uncertainty [43.87879175532034]
近年の論文では、LLMは因果関係分類においてほぼランダムな精度を達成している。
因果的事例への事前曝露が因果的理解を改善するか否かを検討する。
論文 参考訳(メタデータ) (2025-09-24T13:06:35Z) - Benchmarking Reasoning Robustness in Large Language Models [76.79744000300363]
新規データや不完全データでは,性能が著しく低下することがわかった。
これらの結果は、厳密な論理的推論に対するリコールへの依存を浮き彫りにした。
本稿では,情報不足によって引き起こされる幻覚を利用して推論ギャップを明らかにする,Math-RoBと呼ばれる新しいベンチマークを提案する。
論文 参考訳(メタデータ) (2025-03-06T15:36:06Z) - Chest x-ray automated triage: a semiologic approach designed for
clinical implementation, exploiting different types of labels through a
combination of four Deep Learning architectures [83.48996461770017]
本研究では,異なる畳み込みアーキテクチャの後期融合に基づく深層学習手法を提案する。
公開胸部x線画像と機関アーカイブを組み合わせたトレーニングデータセットを4つ構築した。
4つの異なるディープラーニングアーキテクチャをトレーニングし、それらのアウトプットとレイトフュージョン戦略を組み合わせることで、統一されたツールを得ました。
論文 参考訳(メタデータ) (2020-12-23T14:38:35Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。