論文の概要: Kurate: Scalable Scientific Quality Analysis
- arxiv url: http://arxiv.org/abs/2610.07306v1
- Date: Mon, 05 Oct 2026 19:44:07 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-08 02:58:29.619154
- Title: Kurate: Scalable Scientific Quality Analysis
- Title(参考訳): Kurate: スケーラブルな科学的品質分析
- Abstract要約: 本研究では,大規模言語モデル(LLM)を用いた論文の質評価システムKurateを紹介する。
クレートを4,347紙(うちランダム化試験を報告)のコーパスに適用し、各論文を8次元のリサーチデザインとレポートで評価した。
コーパス全体では、論文は統計力、選択的報告、分析事前特定に関する問題が多いことが判明した。
- 参考スコア(独自算出の注目度): 2.395942107244512
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Scientific search systems can find papers that are relevant to a question, but they generally do not assess the quality of the evidence that those papers provide. We present Kurate, a system that uses large language models (LLMs) to assess the quality of published studies. Kurate uses both the paper and its related documents (e.g., the study's trial registration and protocol), and links each of its judgments to the passage of text on which that judgment is based. We applied Kurate to a corpus of 4,347 papers (3,913 of which report randomized trials) and scored each paper on 8 dimensions of study design and reporting: specifically, statistical power, causal identification, preregistration, selective reporting, measurement validity, analysis prespecification, reporting transparency, and conflict of interest and funding. Across the corpus, we found that papers most often exhibited issues with statistical power, selective reporting, and analysis prespecification, although average quality differed between clinical areas. When compared against expert annotations of 60 held-out clinical-trial documents, the information Kurate extracted matched the expert label in 221/242 protocol scorepoints and 294/370 results-publication scorepoints, with AC1 0.94 and 0.81, respectively. Using a well-reputed, high quality clinical trial as a worked example, we show how a single paper's overall grade breaks down into separate judgments, with each linked to specific evidence from the trial's registration, protocol, and published report. Together, these results show that large-scale quality assessment of this kind is feasible, and that it can be used to address meta-scientific research questions.
- Abstract(参考訳): 科学的検索システムは、質問に関連する論文を見つけることができるが、それらの論文が提供する証拠の質を評価することは一般的にない。
本研究では,大規模言語モデル(LLM)を用いて論文の質を評価するシステムであるKurateを紹介する。
Kurateは、論文とその関連文書(例えば、試験登録とプロトコル)の両方を使用し、それぞれの判断を、その判断がベースとなっているテキストのパスにリンクする。
4,347論文 (3,913件) のコーパスに適用し, 統計力, 因果同定, 事前登録, 選択的報告, 分析事前特定, 透明性の報告, 関心と資金の対立など, 研究設計と報告の8つの側面について各論文を採点した。
コーパス全体では,臨床領域間で平均品質が異なっていたが,統計力,選択的報告,分析事前特定などの問題が多い。
臨床検診資料60件の専門家アノテーションと比較すると, クレートが抽出した情報は, 221/242プロトコルスコアポイントと294/370結果スコアポイントで, AC1 0.94点, 0.81点で専門家ラベルと一致した。
報告された高い品質の臨床試験を実例として、単一の論文の全体成績が、試験の登録、議定書、公表された報告から特定の証拠に関連付けられて、どのように別個の判断に分解されるかを示す。
これらの結果から, この種の大規模品質評価が実現可能であり, メタ科学的研究の課題に対処できる可能性が示唆された。
関連論文リスト
- PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress [130.58520199343707]
PaperDoctorは,3つの重要なイノベーションを備えた,サブミッション前のフィードバックのためのエージェントフレームワークである。
まず、全体論的階層的なフレームワークは、書き込み、レイアウト、参照、コード、理論、先行作業、実験を評価します。
第二に、各発見には、観察、文、方程式、コードラインなどの特定の証拠へのポインタ、修正提案が含まれる。
第3に、PaperDoctorはクレームの重要性と計算予算に基づいて実験を選択的に再構築し、再実行する。
論文 参考訳(メタデータ) (2026-09-15T11:08:11Z) - HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research [52.87128503345243]
我々は,Deep Researchにおける階層的エビデンス・アグリゲーションを評価するためのベンチマークであるHiEviDR-Benchを紹介する。
HiEviDR-Benchは、エビデンスの選択、クロスソースリンク、アグリゲーションをキャプチャする明示的なエビデンスグラフを持つ各インスタンスを表す。
報告品質,エビデンストレーサビリティ,引用精度,クレーム検証,回答正当性という5次元のトレーサビリティ指向評価フレームワークを開発した。
論文 参考訳(メタデータ) (2026-07-27T23:50:46Z) - An Audit of Machine Learning Experiments on Software Defect Prediction [1.2743036577573925]
機械学習アルゴリズムは、欠陥のあるソフトウェアコンポーネントを予測するために広く使われている。
本稿では,最近のソフトウェア欠陥予測(SDP)研究を,その設計,解析,報告の実践から評価する。
論文 参考訳(メタデータ) (2026-01-26T13:31:32Z) - Semantic Properties of cosine based bias scores for word embeddings [48.0753688775574]
本稿では,バイアスの定量化に有効なバイアススコアの要件を提案する。
これらの要件について,コサインに基づくスコアを文献から分析する。
これらの結果は、バイアススコアの制限がアプリケーションケースに影響を及ぼすことを示す実験で裏付けられている。
論文 参考訳(メタデータ) (2024-01-27T20:31:10Z) - CausalCite: A Causal Formulation of Paper Citations [80.82622421055734]
CausalCiteは紙の意義を測定するための新しい方法だ。
これは、従来のマッチングフレームワークを高次元のテキスト埋め込みに適応させる、新しい因果推論手法であるTextMatchに基づいている。
科学専門家が報告した紙衝撃と高い相関性など,各種基準におけるCausalCiteの有効性を実証する。
論文 参考訳(メタデータ) (2023-11-05T23:09:39Z) - Chain-of-Factors Paper-Reviewer Matching [32.86512592730291]
本稿では,意味的・話題的・引用的要因を協調的に考慮した,論文レビューアマッチングのための統一モデルを提案する。
提案したChain-of-Factorsモデルの有効性を,最先端のペーパー-リビューアマッチング手法と科学的事前学習言語モデルと比較した。
論文 参考訳(メタデータ) (2023-10-23T01:29:18Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。