論文の概要: Authority Signals in Claude AI Health Citations: A Descriptive Analysis Using the Authority Signals Framework
- arxiv url: http://arxiv.org/abs/2605.23921v1
- Date: Fri, 17 Apr 2026 16:32:39 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-15 07:09:36.467687
- Title: Authority Signals in Claude AI Health Citations: A Descriptive Analysis Using the Authority Signals Framework
- Title(参考訳): AIヘルスクレードにおけるオーソリティシグナル:オーソリティシグナルフレームワークを用いた説明的分析
- Authors: Erin T. Jacques, Erela Datuowei, Elizabeth Quaye, Corey H. Basch, Arijit Chatterjee, Juanita Davis,
- Abstract要約: この調査は、Google Researchが収集した3,172の消費者健康問題を含むHealthSearchQAのデータを使用した。
全引用の97.8%を 機関が占めている
医療機関は最も多く引用される組織タイプ(36.5%)で、続いて政府資源(31.6%)と専門協会(28.4%)が続いた。
上位10団体は全引用の57.8%を占め、メイヨー・クリニックは24.7%だった。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by-nc-nd/4.0/
- Abstract: This study seeks to determine the authority signals used by Anthropic's Claude AI in its presentation of sources when answering consumer health questions. While there exists a great deal of discourse around the quality of health citations that LLMs produce, there is limited information on the integrity of the sources the citations originate from, and to what extent the sources are, from what health professionals would consider, credible sources. This descriptive cross-sectional study used data from HealthSearchQA, which contains 3,172 consumer health questions curated by Google Research. After exclusions, a final dataset of 3,075 questions yielding 10,038 citations was analyzed. The Authority Signals Framework (Jacques et al., 2026) was applied to examine 10 authority signals across four domains for a disproportionate stratified sample of 542 sources. Established institutional sources accounted for 97.8% of all citations (n = 9,818). Medical Institutions were the most frequently cited organization type (36.5%), followed by Government Resources (31.6%) and Professional Associations (28.4%). Commercial Health Information comprised 2.2% (n = 220). The top 10 organizations accounted for 57.8% of all citations, with Mayo Clinic alone representing 24.7%. Among commercial sources in the focused sample, 86.4% displayed medical review statements, 82.5% used schema markup, and 71.8% had comprehensive content, while traditional institutional sources appeared in Claude's citations with or without these same markers. As Anthropic positions Claude for HIPAA-ready healthcare applications, these findings establish a baseline for Claude's citation behavior and demonstrate the utility of the Authority Signals Framework as a tool for ongoing, cross-platform evaluation of AI-mediated health information.
- Abstract(参考訳): この研究は、消費者の健康問題に答える際の情報源の提示において、ArthropicのClaude AIが使用する権威信号を決定することを目的としている。
LLMが生み出す健康的な引用の質について、多くの議論があるが、その引用が生み出す情報源の完全性や、その情報源がどの程度まで、医療専門家が考慮すべき信頼できる情報源であるかについては、限られた情報がある。
この説明的横断的研究は、Google Researchが収集した3,172の消費者健康問題を含むHealthSearchQAのデータを使用した。
除外後,3,075問の最終データセットから10,038問を抽出した。
The Authority Signals Framework (Jacques et al , 2026) was applied to examine 10 authority signal across four domain for a disproportionate stratified sample of 542 sources。
全引用の97.8%(n = 9,818)の機関資料が確立した。
医療機関は最も多く引用される組織タイプ(36.5%)で、続いて政府資源(31.6%)と専門協会(28.4%)が続いた。
商業健康情報は2.2%(n = 220)であった。
上位10団体は全引用の57.8%を占め、メイヨー・クリニックは24.7%だった。
焦点を絞ったサンプルでは86.4%が医療査定文、82.5%がスキーママークアップ、71.8%が包括的コンテンツであり、一方で伝統的な機関資料はクロードの引用に同じマーカーの有無で表示されていた。
HIPAA対応ヘルスケアアプリケーションのためのClaudeの位置づけとして、これらの発見は、Claudeの引用行動のベースラインを確立し、AIを介する健康情報の継続的なクロスプラットフォームな評価ツールとしての Authority Signals Frameworkの有用性を実証する。
関連論文リスト
- Measuring Epistemic Resilience of LLMs Under Misleading Medical Context [47.20527783144983]
大規模言語モデル(LLMs)は、医療ライセンス試験のエキスパートレベルスコアに到達した。
誤解を招く文脈が LLM が元来正しく答える質問に注入されると、彼らは正しい答えを放棄する。
本研究は, 逆行性てんかんのレジリエンス下での正しい判断を維持できる能力と, 測定にMedMisBenchを導入している。
論文 参考訳(メタデータ) (2026-06-10T16:27:26Z) - Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs [8.721564756242431]
CITETRACEは、ユーザクエリから取得したソースから生成された回答まで、完全な引用チェーンをトレースするデータセットである。
我々は,意図的アライメント,ソース適合性,回答ソースの忠実度に基づいて各引用をスコアする3次元評価フレームワークを設計する。
プールの向こう側では、引用の30.6%がソースを歪めており、27.1%はドメイン不適切なソースから来ている。
論文 参考訳(メタデータ) (2026-05-27T14:54:05Z) - The Provenance Gap in Clinical AI: Evidence-Traceable Temporal Knowledge Graphs for Rare Disease Reasoning [1.4676581933580473]
最前線の大規模言語モデルは臨床的に正確な出力を生成するが、それらの引用は製造されている。
HEG-TKGは,4つのPubMedレコードと高品質な階層化と1,280の病的軌跡を持つキュレートされたソースから構築された時間的知識グラフに臨床的主張を基礎づけるシステムである。
論文 参考訳(メタデータ) (2026-04-18T19:10:12Z) - A Decade-Scale Benchmark Evaluating LLMs' Clinical Practice Guidelines Detection and Adherence in Multi-turn Conversations [60.2076951536797]
大規模言語モデル(LLM)は、医療シナリオにますます多くデプロイされている。
LLMが会話中に臨床ガイドラインを特定・遵守できるのかは不明確である。
CPGBenchは、LSMの臨床ガイドラインの検出と付着能力をベンチマークする自動フレームワークである。
論文 参考訳(メタデータ) (2026-03-26T09:00:55Z) - SourceBench: Can AI Answers Reference Quality Web Sources? [14.668125843739423]
SourceBenchは、100の現実世界のクエリで引用されたWebソースの品質を測定するためのベンチマークである。
我々は8つの大言語モデル(LLM)、Google検索、および3つのAI検索ツールを、SourceBenchを用いて3996以上の引用ソースで評価した。
論文 参考訳(メタデータ) (2026-02-18T23:15:32Z) - GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models [22.147294042024836]
キュテーションは科学的主張を信頼する基盤を提供するが、それらが無効または製造された場合、この信頼は崩壊する。
LLM(Large Language Models)の出現により、このリスクは増大した。
我々は大規模な引用検証のためのオープンソースのフレームワークであるCiteVerifierを開発した。
論文 参考訳(メタデータ) (2026-02-06T14:08:34Z) - Authority Signals in AI Cited Health Sources: A Framework for Evaluating Source Credibility in ChatGPT Responses [0.0]
この調査は、Google Researchが収集した3,173の消費者健康質問を含むHealthSearchQAから、ランダムに100の質問を選択した。
これらの質問はChatGPT 5.2 Proに入力され、オーソリティ・シグナル・フレームワークの4つのドメインのレンズを通して引用されたソースを記録し、コーディングした。
ChatGPTの健康に起因した回答の75%以上は、メイヨー・クリニック、クリーブランド・クリニック、ウィキペディア、ナショナル・ヘルス・サービスなど、確立された機関からのものだった。
論文 参考訳(メタデータ) (2026-01-23T17:44:36Z) - An Agentic System for Rare Disease Diagnosis with Traceable Reasoning [69.46279475491164]
大型言語モデル(LLM)を用いた最初のまれな疾患診断エージェントシステムであるDeepRareを紹介する。
DeepRareは、まれな疾患の診断仮説を分類し、それぞれに透明な推論の連鎖が伴う。
このシステムは2,919の疾患に対して異常な診断性能を示し、1013の疾患に対して100%の精度を達成している。
論文 参考訳(メタデータ) (2025-06-25T13:42:26Z) - MedAlign: A Clinician-Generated Dataset for Instruction Following with
Electronic Medical Records [60.35217378132709]
大型言語モデル(LLM)は、人間レベルの流布で自然言語の指示に従うことができる。
医療のための現実的なテキスト生成タスクにおけるLCMの評価は依然として困難である。
我々は、EHRデータのための983の自然言語命令のベンチマークデータセットであるMedAlignを紹介する。
論文 参考訳(メタデータ) (2023-08-27T12:24:39Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。