論文の概要: What Makes a Good AI Review? Concern-Level Diagnostics for AI Peer Review
- arxiv url: http://arxiv.org/abs/2604.19998v1
- Date: Tue, 21 Apr 2026 21:16:59 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-04-23 15:36:10.848046
- Title: What Makes a Good AI Review? Concern-Level Diagnostics for AI Peer Review
- Title(参考訳): 良いAIレビューとは何か? AIピアレビューの懸念レベル診断
- Authors: Ming Jin,
- Abstract要約: 本稿では,AIレビューを判断レベルでのみ評価するのではなく,関心レベルで評価する診断フレームワークを提案する。
本稿では,二項精度から問題検出,判定・階層化動作,判断・認識の校正,帰属・認識の分解に移行した評価ラグを導出する。
- 参考スコア(独自算出の注目度): 6.59569431190247
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Evaluating AI-generated reviews by verdict agreement is widely recognized as insufficient, yet current alternatives rarely audit which concerns a system identifies, how it prioritizes them, or whether those priorities align with the review rationale that shaped the final assessment. We propose concern alignment, a diagnostic framework that evaluates AI reviews at the concern level rather than only at the verdict level. The framework's core data structure is the match graph, a bipartite alignment between official and AI-generated concerns annotated with match type, severity, and post-rebuttal treatment. From this artifact we derive an evaluation ladder that moves from binary accuracy to concern detection, verdict-stratified behavior, decision-aware calibration, and rebuttal-aware decomposition. In a pilot study of four public AI review systems evaluated in six configurations, concern-level analysis suggests that detection alone does not determine review quality; calibration is often the binding constraint. Systems detect non-trivial fractions of official concerns yet most mark 25--55% of concerns on accepted papers as decisive, where, under our operationalization, no official concern on accepted papers was treated as a decisive blocker. Identical overall verdict accuracy can conceal reject-heavy behavior versus low-recall profiles, and low full-review false decisive rates can partly reflect concern dilution rather than calibrated prioritization. Most systems do not emit a native accept/reject, and inferring it from review tone is method-sensitive, reinforcing the need for concern-level diagnostics that remain stable across inference choices. The contribution is a reusable evaluation framework for auditing which concerns AI reviewers identify, how they weight them, and whether those priorities align with the review rationale that informed the paper's final assessment.
- Abstract(参考訳): 判断合意によるAI生成レビューの評価は不十分であると広く認識されているが、現在の代替案では、システムが特定する関心事や優先順位、あるいは最終的な評価を形作るレビューの根拠に合致するかどうかを検査することは滅多にない。
我々は、AIレビューを判断レベルでのみ評価するのではなく、関心レベルで評価する診断フレームワークである、懸念アライメントを提案する。
フレームワークの中核となるデータ構造は、マッチグラフ(Match graph)である。
このアーティファクトから、二分精度から問題検出、判定・階層化行動、判断・認識の校正、帰属・認識の分解に移行した評価ラグを導出する。
6つの構成で評価された4つの公開AIレビューシステムのパイロットスタディにおいて、関心レベル分析は、検出のみがレビューの品質を判断しないことを示唆している。
システムは、非自明な公式な懸念を検知するが、ほとんどの場合、受理された論文に対する懸念の25~55%が決定的であり、当社の運用下では、受理された論文に対する公式な懸念は決定的妨害として扱われなかった。
確定的な全体的な判定精度は、低リコールプロファイルに対する拒絶重みの挙動を隠蔽し、フルリビューの偽判定率を低くすることは、優先順位付けを校正するよりも懸念の希釈を部分的に反映することができる。
ほとんどのシステムはネイティブのアセプション/リジェクトを出力しておらず、レビュートーンからそれを推測することはメソッドに敏感であり、推論選択全体で安定している関心レベル診断の必要性を補強する。
このコントリビューションは、AIレビュアーが識別し、どのように重み付けするか、そしてそれらの優先順位が、論文の最終評価を知らせるレビューの根拠に一致しているかを懸念する監査のための再利用可能な評価フレームワークである。
関連論文リスト
- Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews [69.66583722746904]
私たちは、AIレビュアーを5次元にわたって評価する総合的な評価フレームワークであるBeyond Ratingを紹介します。
本稿では,専門家の不一致に対応するためのMax-Recall戦略を提案する。
提案したテキスト中心の指標は、特に弱みの議論のリコールであり、評価精度と強く相関している。
論文 参考訳(メタデータ) (2026-04-21T14:21:15Z) - PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review [54.141490756509306]
本稿では、エラーデータセットであるPaperAudit-Datasetと、自動レビューフレームワークであるPaperAudit-Reviewの2つのコンポーネントからなるPaperAudit-Benchを紹介する。
PaperAudit-Benchの実験では、モデルと検出深さの誤差検出可能性に大きなばらつきが示された。
本研究では,SFTおよびRLによる軽量LLM検出器のトレーニングをサポートし,計算コストの削減による効率的な誤り検出を実現する。
論文 参考訳(メタデータ) (2026-01-07T04:26:12Z) - TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them [58.04324690859212]
自動評価器(LLM-as-a-judge)としての大規模言語モデル(LLM)は、現在の評価フレームワークにおいて重大な矛盾を明らかにしている。
スコア比較不整合とペアワイズ・トランジティビティ不整合という2つの基本的不整合を同定する。
我々は2つの重要なイノベーションを通じてこれらの制限に対処する確率的フレームワークであるTrustJudgeを提案する。
論文 参考訳(メタデータ) (2025-09-25T13:04:29Z) - How to Evaluate Medical AI [4.23552814358972]
アルゴリズム診断(RPAD, RRAD)の相対精度とリコールについて紹介する。
RPADとRADは、AIの出力を単一の参照ではなく複数の専門家の意見と比較する。
大規模な研究によると、DeepSeek-V3のようなトップパフォーマンスモデルは、専門家のコンセンサスに匹敵する一貫性を達成している。
論文 参考訳(メタデータ) (2025-09-15T14:01:22Z) - Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework [55.078301794183496]
我々は、高品質なピアレビューを支えるコアレビュースキル、すなわち欠陥のある研究ロジックの検出に注力する。
これは、論文の結果、解釈、クレームの間の内部の一貫性を評価することを含む。
本稿では,このスキルを制御条件下で分離し,テストする,完全自動対物評価フレームワークを提案する。
論文 参考訳(メタデータ) (2025-08-29T08:48:00Z) - Evaluating AI systems under uncertain ground truth: a case study in dermatology [43.8328264420381]
不確実性を無視することは、モデル性能の過度に楽観的な推定につながることを示す。
皮膚状態の分類では,データセットの大部分が重大な真理不確実性を示すことが判明した。
論文 参考訳(メタデータ) (2023-07-05T10:33:45Z) - Consultation Checklists: Standardising the Human Evaluation of Medical
Note Generation [58.54483567073125]
本稿では,コンサルテーションチェックリストの評価を基礎として,客観性向上を目的としたプロトコルを提案する。
このプロトコルを用いた最初の評価研究において,アノテータ間合意の良好なレベルを観察した。
論文 参考訳(メタデータ) (2022-11-17T10:54:28Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。