論文の概要: HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment
- arxiv url: http://arxiv.org/abs/2608.10584v2
- Date: Wed, 12 Aug 2026 09:09:14 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-13 14:28:20.606573
- Title: HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment
- Title(参考訳): HexEval: 多次元スカラーアセスメントのためのエビデンス駆動ヘキサゴナルフレームワーク
- Abstract要約: HexEvalは、多次元の学術評価のためのエビデンス駆動の六角形フレームワークである。
HexEvalは評価過程を通じて中間的証拠,次元特異な有理性,検証信号を保持する。
- 参考スコア(独自算出の注目度): 4.103608919706051
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Scholar assessment plays a fundamental role in faculty recruitment, funding allocation, academic promotion, and talent discovery. Existing scholar assessment methods predominantly rely on bibliometric indicators and reputation proxies, while recent large language model (LLM)-based approaches mainly focus on evaluating individual research papers rather than comprehensively assessing scholars. We argue that scholar assessment should be formulated as an evidence-driven reasoning problem that jointly considers intrinsic research quality and externally verifiable scholarly behavior. To this end, we propose HexEval, an evidence-driven hexagonal framework for multidimensional scholar assessment. HexEval explicitly organizes scholar assessment into two complementary evidence layers. The intrinsic layer evaluates anonymized representative works along three dimensions, namely research rigor, methodological innovation, and scientific contribution, whereas the external layer characterizes scholars through knowledge translation, research coherence, and academic impact using heterogeneous evidence collected from GitHub, Lens, OpenAlex, and other publicly verifiable sources. Instead of producing opaque aggregate scores, HexEval preserves intermediate evidence, dimension-specific rationales, and verification signals throughout the evaluation process, enabling interpretable and auditable scholar profiles. Experiments across all six dimensions show dimension-dependent agreement with human or external reference criteria: structured calibration improves absolute agreement for intrinsic quality, while the external modules recover broad trajectory and ordinal impact signals. These results support evidence-driven reasoning over heterogeneous scholarly evidence as a promising paradigm for auditable AI-assisted scholar assessment, while exposing the coverage and attribution limitations of public scholarly data.
- Abstract(参考訳): 学生評価は、教員の募集、資金配分、学術的昇進、人材発見において基本的な役割を担っている。
既存の学者評価手法は、主に文献指標と評価プロキシに依存しているが、最近の大規模言語モデル(LLM)に基づくアプローチは、学者を包括的に評価するのではなく、個々の研究論文を評価することに重点を置いている。
学術的評価は,本質的な研究品質と外的検証可能な学術行動とを共同で考慮した証拠駆動推論問題として定式化されるべきである。
この目的のために,多次元学術評価のためのエビデンス駆動ヘキサゴナルフレームワークであるHexEvalを提案する。
HexEvalは、学者評価を2つの補完的な証拠層に明確に整理している。
内在的な層は、研究厳密性、方法論的革新、科学的貢献という3つの側面に沿って匿名化された代表作品を評価する一方、外部層は、GitHub、Lens、OpenAlex、その他の公的に検証された情報源から収集された異質な証拠を用いて、知識翻訳、研究コヒーレンス、学術的影響を通じて学者を特徴づける。
HexEvalは不透明なアグリゲーションスコアを生成する代わりに、評価プロセス全体を通して中間的証拠、次元固有の有理性、検証信号を保持し、解釈可能で監査可能な学者プロファイルを可能にする。
構造的キャリブレーションは内在的品質の絶対的一致を改善し、外部モジュールは広い軌道と順序の衝撃信号を回復する。
これらの結果は、異質な学術的証拠に対するエビデンス駆動推論を、公的な学術的データのカバレッジと帰属の限界を露呈しつつ、監査可能なAI支援学術的評価のための有望なパラダイムとして支持する。
関連論文リスト
- Ontology-Guided Multi-Agent Extraction of Evaluation Objects from Academic Review Texts: Evidence from Chinese Library and Information Science [6.306930443447107]
本研究では,オブジェクト抽出評価のためのオントロジー誘導型マルチエージェントフレームワークを提案する。
90.33%の精度、84.55%のリコール、87.34%のエンティティレベルF1、79.78%のStrict Typed F1、91.35%の型精度を実現している。
論文 参考訳(メタデータ) (2026-08-30T03:09:47Z) - PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs [68.27437550335709]
本稿では,研究論文に対する統合的およびエージェント指向の科学的推論を評価するためのベンチマークであるPaperMindを紹介する。
PaperMindは、農業、生物学、化学、計算機科学、医学、物理学、経済学を含む7つの領域にわたる実際の科学論文から構築されている。
複数のタスクにわたるモデル行動を分析することにより、PaperMindは統合された科学的推論行動の診断的評価を可能にする。
論文 参考訳(メタデータ) (2026-04-23T05:42:39Z) - InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem [87.30601926271864]
InnoEvalは、人間レベルのアイデアアセスメントをエミュレートするために設計された、深いイノベーション評価フレームワークである。
我々は,多様なオンライン情報源から動的証拠を検索し,根拠とする異種深層知識検索エンジンを適用した。
InnoEvalをベンチマークするために、権威あるピアレビューされた提案から派生した包括的なデータセットを構築します。
論文 参考訳(メタデータ) (2026-02-16T00:40:31Z) - The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research [56.80927148740585]
我々は、動的に進化し、研究評価者としてAIエージェントを開発することで、スケーラビリティと厳密さの課題に対処する。
我々は,機械的解釈可能性の研究をテストベッドとして使用し,標準化された研究成果を構築し,MechEvalAgentを開発した。
我々の研究は、AIエージェントが研究評価を変革し、厳格な科学的実践の道を開く可能性を実証している。
論文 参考訳(メタデータ) (2026-02-05T19:00:02Z) - ScholarPeer: A Context-Aware Multi-Agent Framework for Automated Peer Review [48.60540055009675]
ScholarPeerは、上級研究者の認知過程をエミュレートするために設計された、検索可能なマルチエージェントフレームワークである。
We evaluate ScholarPeer on DeepReview-13K and the results showed that ScholarPeer achieve significant win-rates against state-of-the-art approach in side-side-side evaluations。
論文 参考訳(メタデータ) (2026-01-30T06:54:55Z) - DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Survey [53.85391477976017]
DeepSurvey-Benchは、生成された調査の学術的価値を包括的に評価するために設計された、新しいベンチマークである。
学術的価値アノテーションを用いた信頼性のあるデータセットを構築し, 生成した調査の深い学術的価値を評価する。
論文 参考訳(メタデータ) (2026-01-13T14:42:56Z) - ScholarEval: Research Idea Evaluation Grounded in Literature [18.31628500009905]
ScholarEvalは2つの基本的な基準に基づいて研究アイデアを評価する検索強化評価フレームワークである。
ScholarEvalを評価するために、ScholarIdeasを紹介します。
以上の結果から,ScholarEvalは,ScholarIdeasのアノテートルーリックに言及される点を,すべての基線に比べてはるかに高い範囲でカバーできることが示唆された。
論文 参考訳(メタデータ) (2025-10-17T21:55:07Z) - Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework [55.078301794183496]
我々は、高品質なピアレビューを支えるコアレビュースキル、すなわち欠陥のある研究ロジックの検出に注力する。
これは、論文の結果、解釈、クレームの間の内部の一貫性を評価することを含む。
本稿では,このスキルを制御条件下で分離し,テストする,完全自動対物評価フレームワークを提案する。
論文 参考訳(メタデータ) (2025-08-29T08:48:00Z) - A Content-Based Novelty Measure for Scholarly Publications: A Proof of
Concept [9.148691357200216]
学術出版物にノベルティの情報理論尺度を導入する。
この尺度は、学術談話の単語分布を表す言語モデルによって知覚される「サプライズ」の度合いを定量化する。
論文 参考訳(メタデータ) (2024-01-08T03:14:24Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。