論文の概要: OmniMed-Jev: Calibrating LVLM Confidence for Trustworthy Medical Multimodal Decisions via System One
- arxiv url: http://arxiv.org/abs/2610.00381v1
- Date: Wed, 30 Sep 2026 08:53:54 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-03 01:19:23.627881
- Title: OmniMed-Jev: Calibrating LVLM Confidence for Trustworthy Medical Multimodal Decisions via System One
- Title(参考訳): OmniMed-Jev: System Oneによる信頼できる医療マルチモーダル決定のためのLVLM信頼性の校正
- Abstract要約: 提案するOmniMed-Jevは,複数モーダル供給候補集合に対する選択,Nuul,スコア決定として,それぞれの医学的決定を表現している。
様々な画像モダリティを受け入れ、異なる予測タスクをカバーし、1つの候補条件付き確率モデルを通してそれらを表現する。
OmniMed-Jevの報告した確率トラックは、同じバックボーン、データ、スケジュールでトレーニングされた生成ベースラインに対するインターフェース制御された比較において、より正確に観察された。
- 参考スコア(独自算出の注目度): 9.555285052728198
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Medical models are judged not only on correctness, but on whether reported confidence matches actual accuracy. Generalist multimodal medical models have expanded what a single model can perceive, yet they still express bounded decisions such as diagnoses, findings or cell counts as generated text, so the reported probability reflects the next token rather than the decision itself. Motivated by decision-native interfaces such as Jev, we introduce OmniMed-Jev, which represents each medical decision as a Choice, Noul or Score decision over a runtime-supplied candidate set and returns a full distribution over that set: mutually exclusive classes, binary presence of a finding, or a bounded ordered value. The design is omni in three respects: it accepts diverse imaging modalities, covers different prediction tasks, and expresses them through one candidate-conditioned probability model, so heterogeneous outputs become comparable probabilities rather than task-specific strings. In an interface-controlled comparison against a generative baseline trained on the same backbone, data and schedule, OmniMed-Jev's reported probabilities track observed correctness far more closely, reducing calibration error by up to an order of magnitude and reliability error by up to two, while point-prediction performance remains comparable; counting is the one family where the generative baseline stays ahead. Making the decision distribution the model's output is not a format change but what turns reported numbers into probabilities that mean what they say. These results support explicit decision modeling as a way to make reported confidence meaningful within the evaluated tasks, and they are not evidence of clinical readiness: the comparison cannot separate the interface from associated training differences, which we state alongside the results. Code is available at github.com/lytang63/OmniMed-Jev.
- Abstract(参考訳): 医療モデルは、正確性だけでなく、報告された信頼度が実際の正確性に合致するかどうかに基づいて判断される。
一般的なマルチモーダル医療モデルは、単一のモデルが知覚できるものを拡張してきたが、診断、発見、細胞数などの有界な決定を生成テキストとして表現しているため、報告された確率は、決定そのものよりも次のトークンを反映している。
Jevのような決定ネイティブインターフェースによって動機づけられたOmniMed-Jevを紹介します。これは、実行時に供給される候補セットに対して、それぞれの医学的決定をチョイス、ヌール、スコア決定として表現し、そのセット上の完全な分布を返します。
多様な画像モダリティを受け入れ、異なる予測タスクをカバーし、1つの候補条件付き確率モデルを通してそれらを表現することにより、不均一な出力はタスク固有の文字列よりも同等の確率となる。
同じバックボーン、データ、スケジュールでトレーニングされた生成ベースラインに対するインタフェース制御比較において、OmniMed-Jevの報告された確率トラックは、はるかに正確に観測され、キャリブレーション誤差を最大2桁まで低減し、ポイント予測性能は同等であり、生成ベースラインが先行する1つのファミリーである。
決定を分散させることで、モデルの出力はフォーマットの変更ではなく、レポートされた数値を、彼らが何を言っているかを意味する確率に変換する。
これらの結果は,評価課題において,報告された信頼度を有意に評価する方法として,明確な意思決定モデリングを支持している。
コードはgithub.com/lytang63/OmniMed-Jevで入手できる。
関連論文リスト
- Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding [65.85599131866623]
私たちはJevとContractNLIの9つの言語モデルを比較します。
制御された比較は、仮説の可視性、要求された出力、および出力順序によって異なる。
Jevは、評価された構成の中で、最低コストと中央値のレスポンス時間を持っています。
論文 参考訳(メタデータ) (2026-09-23T10:53:01Z) - DistMedVL: Distributional Vision-Language Alignment for Uncertainty-Aware Medical Image Segmentation [16.13084489571617]
視覚的およびテキスト的表現の相互アライメントは、マルチモーダルな医用画像理解に不可欠である。
既存の視覚言語セグメンテーション法は、不明瞭な境界からアレター的不確実性を見落としている決定論的クロスモーダルマッチングに依存している。
本稿では,凍結エンコーダ上に軽量な確率的クロスモーダルアダプタ(PCM-Adapter)を導入し,表現の不確実性を明示的にモデル化する確率的視覚言語フレームワークであるDistMedVLを提案する。
論文 参考訳(メタデータ) (2026-08-06T07:17:14Z) - EnTrust: Modeling Inter-Modal Conflict for Trustworthy Multimodal Medical Image Analysis [41.67522319355107]
EnTrustは、モーダル間紛争を予測の不確実性の主要な原因として扱うフレームワークである。
最強のベースラインに比べてキャリブレーション誤差を40%削減しつつ、最先端のセグメンテーション精度を実現する。
論文 参考訳(メタデータ) (2026-06-19T12:43:50Z) - Internal Representation, Not Clinical Knowledge: Where Apparent LLM Triage Failures Originate [2.581200752140087]
患者投票型臨床トリアージベンチマークでは、制約付き多重選択出力における消費者LCMの低トライアージ率が高いことが報告されている。
両フォーマットで共有された臨床物語に同一の医療的特徴が現れるが、すべてのモデルにおいて、複数の選択決定トークンに沈黙する。
論文 参考訳(メタデータ) (2026-05-28T13:14:17Z) - MedObvious: Exposing the Medical Moravec's Paradox in VLMs via Clinical Triage [20.835664121303534]
ビジョン言語モデル(VLM)は、医療報告生成や視覚的質問応答といったタスクにますます使われています。
臨床実践では、解釈は診断前の衛生検査から始まる。
既存のベンチマークでは、このステップが解決されたと仮定しており、致命的な障害モードを見逃している。
我々は1,880タスクのベンチマークであるMedObviousを導入し、入力検証をセットレベルの一貫性機能として分離する。
論文 参考訳(メタデータ) (2026-03-24T17:59:54Z) - Multidimensional Uncertainty Quantification via Optimal Transport [87.97146725546502]
相補的なUQ測度をベクトルに積み重ねることで,不確実性定量化(UQ)の多次元的考察を行う。
VecUQ-OTは、個々の測定が失敗しても高い効率を示す。
論文 参考訳(メタデータ) (2025-09-26T14:09:03Z) - Uncertainty Estimation of Large Language Models in Medical Question Answering [60.72223137560633]
大規模言語モデル(LLM)は、医療における自然言語生成の約束を示すが、事実的に誤った情報を幻覚させるリスクがある。
医学的問合せデータセットのモデルサイズが異なる人気不確実性推定(UE)手法をベンチマークする。
以上の結果から,本領域における現在のアプローチは,医療応用におけるUEの課題を浮き彫りにしている。
論文 参考訳(メタデータ) (2024-07-11T16:51:33Z) - Towards Reliable Medical Image Segmentation by Modeling Evidential Calibrated Uncertainty [57.023423137202485]
医用画像のセグメンテーションの信頼性に関する懸念が臨床医の間で続いている。
本稿では,医療画像セグメンテーションネットワークにシームレスに統合可能な,実装が容易な基礎モデルであるDEviSを紹介する。
主観的論理理論を活用することで、医用画像分割の確率と不確実性を明示的にモデル化する。
論文 参考訳(メタデータ) (2023-01-01T05:02:46Z) - Statistical inference of the inter-sample Dice distribution for
discriminative CNN brain lesion segmentation models [0.0]
識別畳み込みニューラルネットワーク(CNN)は多くの脳病変セグメンテーションタスクでよく機能している。
識別的CNNのセグメンテーションサンプリングは、訓練されたモデルの堅牢性を評価するために使用される。
厳格な信頼に基づく決定ルールが提案され、特定の患者に対してCNNモデルを拒絶するか受け入れるかが決定される。
論文 参考訳(メタデータ) (2020-12-04T18:18:24Z) - Decision-Making with Auto-Encoding Variational Bayes [71.44735417472043]
変分分布とは異なる後部近似を用いて意思決定を行うことが示唆された。
これらの理論的な結果から,最適モデルに関するいくつかの近似的提案を学習することを提案する。
おもちゃの例に加えて,単細胞RNAシークエンシングのケーススタディも紹介する。
論文 参考訳(メタデータ) (2020-02-17T19:23:36Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。