論文の概要: How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks
- arxiv url: http://arxiv.org/abs/2609.13009v1
- Date: Fri, 11 Sep 2026 16:06:50 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-15 07:38:31.801883
- Title: How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks
- Title(参考訳): 物理分野でのフロンティアモデルはどの程度優れているか?専門家のリグレーディング調査で評価が損なわれ、ベンチマークのほぼ飽和状態が判明
- Abstract要約: 主要な物理学ベンチマークに関する低報告のスコアは、フロンティア言語モデルが高度な物理学に苦しむことを示唆している。
我々は,6つの広く使用されている物理ベンチマークでフロンティアモデルを評価し,専門家と評価することで,これらの知見を再考する。
- 参考スコア(独自算出の注目度): 60.853503196898906
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Low reported scores on leading physics benchmarks, including those featured in the Artificial Analysis Intelligence Index (2026), suggest that frontier language models still struggle with advanced physics, a demanding test of their scientific reasoning and quantitative problem-solving abilities. Yet this impression does not always align with domain experts' experiences using these models in their work. We revisit these reported findings by evaluating frontier models on six widely used physics benchmarks and auditing them with experts, focusing on text-only problems with verifiable final answers. For each subfield of physics, faculty and graduate researchers with relevant expertise carefully review problem statements, reference solutions, and model responses to distinguish genuine model errors from grader errors, incorrect reference solutions, and ambiguous or underspecified questions. Most audited cases initially evaluated as incorrect reflect these benchmarking issues rather than errors in the models' physics reasoning. We then ask experts to address these benchmarking issues by correcting erroneous reference solutions and repairing or excluding flawed questions. We find that GPT-5.6-Sol's measured mean@4 rises from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark, while its corrected pass@4 reaches 94.4% on the 54 retained CritPt challenges. Corrected scores are computed on the retained evaluation subsets following expert review. Scores on the audited subsets of UGPhysics, PRISM-Physics, and PHYBench also rise substantially after correction. These findings suggest that current benchmarks substantially understate frontier models' ability to solve well-posed physics problems. Near-saturation on these closed-ended tasks highlights the need for more demanding, expert-validated evaluations.
- Abstract(参考訳): 2026年のArtificial Analysis Intelligence Index(英語版)など、主要な物理学ベンチマークのスコアの低さは、フロンティア言語モデルがいまだ高度な物理学に苦戦していることを示唆している。
しかし、この印象はドメインの専門家がこれらのモデルを使った経験と必ずしも一致しない。
筆者らは,6つの広く使用されている物理ベンチマークにおけるフロンティアモデルの評価と専門家による評価を行い,最終回答を検証したテキストのみの問題に焦点をあてた。
物理学のサブフィールドごとに、関連する専門知識を持つ学部と大学院の研究者は、問題文、参照解、モデル応答を慎重にレビューし、真のモデルエラーをグレーダエラー、不正参照解、曖昧または不明確な質問と区別する。
当初、ほとんどの監査済みケースは、モデルの物理推論の誤りよりも、これらのベンチマーク問題を反映していると評価された。
次に、誤った参照ソリューションを修正し、欠陥のある問題を修正したり排除したりすることで、これらのベンチマーク問題に対処するよう専門家に求めます。
GPT-5.6-Sol's measured mean@4 increaseds 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark, and its corrected pass@4 reach 94.4% on the 54 maintained CritPt challenges。
補正されたスコアは、専門家レビューの後、保持された評価サブセットに基づいて計算される。
UGPhysics、PRISM-Physics、PHYBenchの監査されたサブセットのスコアも修正後に大幅に上昇した。
これらの結果は、現在のベンチマークが、適切に配置された物理問題を解くためのフロンティアモデルの能力を大幅に過小評価していることを示唆している。
これらのクローズドなタスクのほぼ飽和は、より要求があり、専門家が検証した評価の必要性を強調します。
関連論文リスト
- NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment [54.4265627247104]
大規模言語モデル(LLM)は、主要なAIカンファレンスでピアレビューでますます使用されている。
既存のベンチマークでは、ノベルティを1つの総合的なスコアとして評価している。
本報告では,詳細なノベルティアセスメント診断のための人手によるベンチマークであるNovGaugeについて述べる。
論文 参考訳(メタデータ) (2026-09-10T08:34:56Z) - SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models [1.9873319610438742]
SciCodeは、言語モデルの科学的コーディング能力の標準尺度である。
人工知能インテリジェンス指数(Artificial Analysis Intelligence Index)の略。
最先端のモデルは、SciCodeが示唆しているよりも科学的コーディングにはるかに熟練している。
論文 参考訳(メタデータ) (2026-08-05T15:45:55Z) - TPS-CalcBench: A Benchmark and Diagnostic Evaluation Framework for LLM Analytical Calculation Competence in Hypersonic Thermal Protection System Engineering [0.8921166277011348]
超音速熱保護システム(TPS)の設計では、不正確な停滞点熱流束や境界層計算が破滅的な設計マージン違反を引き起こす可能性がある。
現在の科学的ベンチマークは抽象数学と基礎物理学のみをテストし、最終回答のみを評価し、工学的推論プロセスを無視し、そのような重大な失敗を検出できない。
TPS-CalcBenchは超音速空力と高温ガス力学における閉形式解析計算のための最初の診断ベンチマークである。
論文 参考訳(メタデータ) (2026-04-20T08:46:49Z) - Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark [49.42250115889234]
本研究では,研究レベルの推論タスクにおいて,大規模言語モデル(LLM)をテストするために設計された最初のベンチマークを示す。
CritPtは71の複合研究課題からなる。
現在最先端のLCMは、孤立したチェックポイントを早期に保証しているが、完全な研究スケールの課題を確実に解決できるには程遠い。
論文 参考訳(メタデータ) (2025-09-30T17:34:03Z) - CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics [71.42168240638462]
CMPhysBenchは、凝縮物質物理学における大規模言語モデルの習熟度を評価するように設計されている。
以上の結果から,最高モデルであるGrok-4でさえ,CMPhysBench上での平均SEEDスコアが36点,精度が28%であった。
論文 参考訳(メタデータ) (2025-08-25T15:32:22Z) - PhyX: Does Your Model Have the "Wits" for Physical Reasoning? [49.083544963243206]
既存のベンチマークでは、物理的な推論という、インテリジェンスの重要な側面を捉えられません。
視覚シナリオにおける物理基底推論のモデルキャパシティを評価するために設計された,最初の大規模ベンチマークであるPhyXを紹介する。
論文 参考訳(メタデータ) (2025-05-21T18:33:50Z) - PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models [33.45006997591683]
PHYBenchは、高校から物理オリンピックの難易度まで、500の物理問題のベンチマークである。
PHYBenchはオリジナルのコンテンツを通じてデータの汚染に対処し、欠陥のあるアイテムを除去するために体系的なキュレーションパイプラインを使用する。
PHYBenchはより多くのトークンを活性化し、推論モデル間のより強力な微分を提供する。
論文 参考訳(メタデータ) (2025-04-22T17:53:29Z) - The Impossible Test: A 2024 Unsolvable Dataset and A Chance for an AGI Quiz [0.0]
我々は、675の根本的な解決不可能な問題に対して不確実性を認識できる大規模言語モデル(LLM)の能力を評価する。
62-68%の精度で得られた最良のモデルは、生物学から哲学、数学まで様々な分野において未知であった。
論文 参考訳(メタデータ) (2024-11-20T04:12:29Z) - OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems [62.06169250463104]
我々はOlympiadレベルのバイリンガル・マルチモーダル・サイエンス・ベンチマークであるOlympiadBenchを紹介し、Olympiadレベルの数学と物理学のコンペティションの8,476の問題を特徴とする。
最も優れたモデルであるGPT-4Vはオリンピアドベンチで平均17.97%を獲得し、物理学ではわずか10.74%である。
GPT-4Vの分析では、幻覚、知識欠失、論理的誤信などの問題が指摘されている。
論文 参考訳(メタデータ) (2024-02-21T18:49:26Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。