論文の概要: MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents
- arxiv url: http://arxiv.org/abs/2607.18999v2
- Date: Sun, 26 Jul 2026 15:35:54 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-28 12:52:43.705677
- Title: MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents
- Title(参考訳): MedDDC-Eval:多施設医療相談エージェントの診断と非結合性評価
- Authors: Guofeng Zhang, Yizeng Quan, Huaiyi Fang, Jianwei Lv, Jinyao Liu, Xunxu Duan, Lening An, Yu Ouyang, Junfeng Wang,
- Abstract要約: MedDDC-Evalは, 既往の症例に比較して, 診断非結合性評価ベッドである。
フリーズされた共有診断読影器は、すべてのポリシーで満たされた履歴に適用される。
診断支援、情報取得、効率を報告している。
- 参考スコア(独自算出の注目度): 4.184241883998287
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Evaluating multi-turn medical consultation agents requires judging the diagnostic support provided by the histories they elicit through interaction. Yet coupled evaluation lets each policy both elicit the history and generate the terminal diagnosis, so a diagnosis score confounds the elicited history with the policy's own terminal diagnosis generator. We introduce MedDDC-Eval, a diagnosis-decoupled evaluation testbed over held-out cases derived from medical records and online consultations. It applies the same frozen shared diagnostic reader to every policy-elicited history, holding terminal diagnosis generation fixed across policies and enabling comparison under the shared diagnostic reader. It reports diagnostic support, information-acquisition coverage, and efficiency. LLM-assisted semantic matching followed by deterministic one-to-one assignment makes the diagnosis-trajectory-efficiency (D/T/E) scores auditable. In a fixed-history audit across eight policies, replacing each policy's own generator with the shared diagnostic reader shifts diagnosis F1 by 2.2-19.0 points and reverses 18% and 36% of pairwise orderings on the Record and Dialogue splits. To examine downstream utility, we use standard Group Relative Policy Optimization (GRPO) with a separate training-time reward that targets the same diagnosis and trajectory dimensions. Relative to its Qwen3-32B initialization, the trained policy gains 9.6 and 4.6 aggregate-score points on the held-out Record and Dialogue splits, respectively, and ablating either feedback signal reduces the aggregate score on both. Together, MedDDC-Eval supports comparison under a shared diagnostic reader and evaluation-informed policy development, while complementing end-to-end evaluation when terminal diagnosis generation is also part of the target capability.
- Abstract(参考訳): マルチターン医療相談エージェントの評価には、相互作用を通じて引き起こす履歴によって提供される診断支援を判断する必要がある。
相互に連携した評価により、各ポリシが履歴を抽出し、端末診断を生成することができるため、診断スコアは、ポリシー独自の端末診断生成装置と一致する。
MedDDC-Evalは,医療記録とオンライン相談から得られた保留症例に対する診断分離評価法である。
フリーズド共有診断読影器は、すべてのポリシー付履歴に適用され、ポリシー間で固定された端末診断生成を保持し、共有診断読影器による比較を可能にする。
診断支援、情報取得、効率を報告している。
LLMによるセマンティックマッチングに続いて決定論的1対1の割り当てにより、診断-軌道効率(D/T/E)スコアが監査可能となる。
8つのポリシーにわたる固定履歴監査では、各ポリシーのジェネレータを共有診断リーダに置き換えることで、診断F1を2.2-19.0ポイントシフトさせ、レコードとダイアログの分割におけるペアの順序の18%と36%を逆転させる。
下流の実用性を検討するために、同じ診断と軌跡次元をターゲットにした訓練時間報酬の標準グループ相対政策最適化(GRPO)を用いる。
Qwen3-32Bの初期化に関連して、トレーニングされたポリシーは、保持されたレコードとダイアログの分割において、それぞれ9.6と4.6のアグリゲーションスコアポイントを獲得し、どちらのフィードバック信号も非難することにより、両者のアグリゲーションスコアが減少する。
MedDDC-Evalは、共通診断読影器と評価インフォームドポリシー開発の下での比較をサポートし、端末診断生成も対象能力の一部である場合のエンドツーエンド評価を補完する。
関連論文リスト
- Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA [4.080704298118247]
マルチモーダルAIは、信頼性と解釈性を維持しながら、視覚とテキストの証拠を組み合わせる必要がある。
我々は、質問応答と説明品質のために、9つの文書化されたシステムにまたがる設計選択を解析する。
この発見は、データ融合、説明可能性、レジリエントな評価に基づいて、信頼できるマルチモーダルヘルスケアAIをサポートする。
論文 参考訳(メタデータ) (2026-07-16T17:38:19Z) - Benchmarking Multi-turn Medical Diagnosis: Hold, Lure, and Self-Correction [72.89352076103889]
大規模言語モデル (LLM) は, 臨床情報がすべて一ターンで提供される場合に, 高い精度で診断を行う。
1,035例からなる高忠実多ターン診断ベンチマークであるMINTを導入する。
診断決定に大きな影響を及ぼす3つの永続的な行動パターンを明らかにする。
論文 参考訳(メタデータ) (2026-04-06T00:23:10Z) - Optimizing Point-of-Care Ultrasound Video Acquisition for Probabilistic Multi-Task Heart Failure Detection [3.5617951813421818]
本稿では、RLエージェントが次のビューを選択して取得を終了するパーソナライズされたデータ取得戦略を提案する。
終了時,診断モデルは大動脈狭窄(AS)重症度と左室放出率(LVEF)を同時予測する
本手法は32%の動画を用いて,AS重度分類とLVEF推定における平均平衡精度(bACC)を77.2%達成し,フルスタディ性能と一致した。
論文 参考訳(メタデータ) (2026-02-14T07:56:21Z) - MedAD-R1: Eliciting Consistent Reasoning in Interpretible Medical Anomaly Detection via Consistency-Reinforced Policy Optimization [46.65200216642429]
我々はMedADの最初の大規模マルチモーダル・マルチセンタベンチマークであるMedAD-38Kを紹介し、構造化された視覚質問応答(VQA)ペアとともに、CoT(Chain-of-Thought)アノテーションを特徴付ける。
提案するモデルであるMedAD-R1は、MedAD-38Kベンチマーク上での最先端(SOTA)性能を実現し、強いベースラインを10%以上上回った。
論文 参考訳(メタデータ) (2026-02-01T07:56:10Z) - Towards Reliable Medical LLMs: Benchmarking and Enhancing Confidence Estimation of Large Language Models in Medical Consultation [97.36081721024728]
本稿では,現実的な医療相談におけるマルチターンインタラクションの信頼性を評価するための最初のベンチマークを提案する。
本ベンチマークでは,3種類の医療データを統合し,診断を行う。
本稿では,エビデンスを基盤とした言語自己評価フレームワークであるMedConfを紹介する。
論文 参考訳(メタデータ) (2026-01-22T04:51:39Z) - MedEinst: Benchmarking the Einstellung Effect in Medical LLMs through Counterfactual Differential Diagnosis [13.241795322837861]
MedEinstは,49の疾患に5,383対の臨床症例を比較検討した。
バイアストラップ速度による感受性の測定-正確な診断制御にもかかわらず、誤診断トラップの確率について検討する。
論文 参考訳(メタデータ) (2026-01-10T17:39:25Z) - Evolving Diagnostic Agents in a Virtual Clinical Environment [75.59389103511559]
本稿では,大規模言語モデル(LLM)を強化学習を用いた診断エージェントとして訓練するためのフレームワークを提案する。
本手法は対話型探索と結果に基づくフィードバックによって診断戦略を取得する。
DiagAgentはDeepSeek-v3やGPT-4oなど、最先端の10のLLMを著しく上回っている。
論文 参考訳(メタデータ) (2025-10-28T17:19:47Z) - Extracting Diagnosis Pathways from Electronic Health Records Using Deep
Reinforcement Learning [2.0191844627740254]
我々は,電子カルテから正しい診断を得るために,行動の最適なシーケンスを学習することを目指している。
この課題に様々な深層強化学習アルゴリズムを適用し、貧血の鑑別診断のために、合成だが現実的なデータセットを実験する。
論文 参考訳(メタデータ) (2023-05-10T16:36:54Z) - Semi-Supervised Variational Reasoning for Medical Dialogue Generation [70.838542865384]
医療対話生成には,患者の状態と医師の行動の2つの重要な特徴がある。
医療対話生成のためのエンドツーエンドの変分推論手法を提案する。
行動分類器と2つの推論検出器から構成される医師政策ネットワークは、拡張推論能力のために提案される。
論文 参考訳(メタデータ) (2021-05-13T04:14:35Z) - Towards Causality-Aware Inferring: A Sequential Discriminative Approach
for Medical Diagnosis [142.90770786804507]
医学診断アシスタント(MDA)は、疾患を識別するための症状を逐次調査する対話型診断エージェントを構築することを目的としている。
この研究は、因果図を利用して、MDAにおけるこれらの重要な問題に対処しようとする。
本稿では,他の記録から知識を引き出すことにより,非記録的調査に効果的に答える確率に基づく患者シミュレータを提案する。
論文 参考訳(メタデータ) (2020-03-14T02:05:54Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。