論文の概要: EduArt: An educational-level benchmark for evaluating art history knowledge in large language models
- arxiv url: http://arxiv.org/abs/2607.02007v1
- Date: Thu, 02 Jul 2026 10:43:06 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-03 19:45:08.793682
- Title: EduArt: An educational-level benchmark for evaluating art history knowledge in large language models
- Title(参考訳): EduArt:大規模言語モデルにおける美術史知識の評価のための教育レベルベンチマーク
- Authors: Gianmarco Spinaci, Lukas Klic, Giovanni Colavizza,
- Abstract要約: 既存のアートに焦点を当てた評価は、合成質問に依存しており、アイテムレベルの特性を報告することは滅多にない。
本稿では,マルチモーダルLLMにおける美術史的知識と視覚的推論のための教育レベルのベンチマークであるEduArtを紹介する。
- 参考スコア(独自算出の注目度): 0.764671395172401
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Large language models now score near ceiling on general benchmarks, but these aggregate measures reveal little about how models behave within single disciplines. Existing art-focused evaluations rely on synthetic questions and rarely report item-level properties. This paper introduces EduArt, an educational-level benchmark for art-historical knowledge and visual reasoning in multimodal LLMs. EduArt comprises 871 human-authored questions from Italian secondary-school exercises and US Advanced Placement Art History exams, spanning two languages and seven formats from multiple choice to in-text word placement and error identification. Twelve models from six provider families were evaluated under a default answer-only condition and a motivation condition requiring written justification, and characterized using Classical Test Theory and a logistic regression isolating the effects of format, language, image presence, and model. The benchmark showed strong psychometric properties (mean discrimination 0.514, 82.3 percent good discriminators), while multiple-choice accuracy saturated near ceiling for six models, showing recognition formats alone cannot distinguish frontier models. Format was a strong independent predictor of accuracy: models exceeding 94 percent on multiple choice fell to 23.9 percent on open completion (Claude Opus 4.6) and 6.2 percent on error identification (Claude Sonnet 4.6). The motivation condition changed accuracy in a predominantly negative, family-dependent direction. These dissociations indicate that art-historical knowledge and the ability to deploy it are distinct capabilities, and that single-format benchmarks overestimate what models can reliably do. Mapping this capability profile is a precondition for responsible use of multimodal LLMs in art-historical scholarship, where tasks demand producing and manipulating content rather than selecting from fixed options.
- Abstract(参考訳): 大規模な言語モデルは、一般的なベンチマークでは天井付近にスコアを付けているが、これらの集計は、モデルが単一の規律の中でどのように振る舞うかをほとんど明らかにしていない。
既存のアートに焦点を当てた評価は、合成質問に依存しており、アイテムレベルの特性を報告することは滅多にない。
本稿では,マルチモーダルLLMにおける美術史的知識と視覚的推論のための教育レベルのベンチマークであるEduArtを紹介する。
EduArtは、イタリアの中等教育演習とアメリカの先進的配置美術史試験から、複数の選択肢からテキスト内の単語配置と誤り識別まで、2つの言語と7つのフォーマットにまたがる871人の人間による質問で構成されている。
6つのプロバイダファミリーの12モデルについて,既定の回答のみ条件とモチベーション条件で評価し,古典的テスト理論とロジスティック回帰を用いて形式,言語,画像存在,モデルの効果を分離した。
このベンチマークでは、強い心理測定特性(平均判別0.514、82.3%の良判別器)が示され、6つのモデルの天井付近に多重選択精度が飽和しており、認識形式だけではフロンティアモデルの識別が不可能であった。
複数の選択肢で94%を超えるモデルは、オープンコンプリート(Claude Opus 4.6)で23.9%、エラー識別(Claude Sonnet 4.6)で6.2%に低下した。
動機づけ条件は、主に負の家族依存方向の精度を変化させた。
これらの解離は、美術史的な知識とそれをデプロイする能力が異なる能力であり、単一フォーマットのベンチマークがモデルが確実にできることを過大評価していることを示している。
この能力プロファイルをマッピングすることは、美術史学におけるマルチモーダル LLM の責任ある使用の前提条件であり、そこでは、特定の選択肢から選択するのではなく、コンテンツの生成と操作をタスクが要求する。
関連論文リスト
- Revisiting Northrop Frye's Four Myths Theory with Large Language Models [0.0]
ノースロップ・フライの4つの基本的物語ジャンルの理論は文学的批判に大きな影響を与えた。
パターンに基づく分析を補完する新しいキャラクタ関数フレームワークを提案する。
4つの普遍的性格関数(主人公,メンター,アンタゴニスト,仲間)をユングの精神構造成分にマッピングすることで導出する。
論文 参考訳(メタデータ) (2026-02-17T16:02:52Z) - AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models [0.0]
AA-Omniscienceは6000の質問に対する事実的リコールと知識のキャリブレーションを測定するために設計されたベンチマークである。
モデルの評価は、事実のリコールを測定する有界メトリック(-100から100)であるOmniscience Indexを測定する。
その結果、フロンティアモデル全体の持続的な事実性とキャリブレーションの弱点が明らかになった。
論文 参考訳(メタデータ) (2025-11-17T06:27:16Z) - Comparative Insights from 12 Machine Learning Models in Extracting Economic Ideology from Political Text [0.0]
本研究では、経済イデオロギーの検出において、12の機械学習モデルとモデルバリエーションの能力を体系的に評価する。
この分析は、粒度および集合レベルでのいくつかの生成、微調整、ゼロショットモデルの性能を評価する。
論文 参考訳(メタデータ) (2025-01-16T18:06:22Z) - Eureka: Evaluating and Understanding Large Foundation Models [23.020996995362104]
Eurekaは、シングルスコアのレポートやランキングを超えて、大規模な基盤モデルの評価を標準化するためのオープンソースのフレームワークです。
我々は、12の最先端モデルを分析し、失敗理解とモデル比較に関する詳細な洞察を提供する。
論文 参考訳(メタデータ) (2024-09-13T18:01:49Z) - ArtGPT-4: Towards Artistic-understanding Large Vision-Language Models with Enhanced Adapter [19.830089364830066]
ArtGPT-4は、芸術的理解における既存のモデルの限界に対処するために設計された大きな視覚言語モデルである。
芸術的理解で画像を描画し、それらが刺激する感情を伝え、人間の解釈を反映する。
論文 参考訳(メタデータ) (2023-05-12T14:04:30Z) - Large Language Models in the Workplace: A Case Study on Prompt
Engineering for Job Type Classification [58.720142291102135]
本研究では,実環境における職種分類の課題について検討する。
目標は、英語の求職が卒業生やエントリーレベルの地位に適切かどうかを判断することである。
論文 参考訳(メタデータ) (2023-03-13T14:09:53Z) - Large Language Models Are Latent Variable Models: Explaining and Finding
Good Demonstrations for In-Context Learning [104.58874584354787]
近年,事前学習型大規模言語モデル (LLM) は,インコンテキスト学習(in-context learning)として知られる推論時少数ショット学習能力を実現する上で,顕著な効率性を示している。
本研究では,現実のLLMを潜在変数モデルとみなし,ベイズレンズによる文脈内学習現象を考察することを目的とする。
論文 参考訳(メタデータ) (2023-01-27T18:59:01Z) - Labeling Explicit Discourse Relations using Pre-trained Language Models [0.0]
最先端のモデルは手作りの機能を使ってFスコアの45%をわずかに上回っている。
事前訓練された言語モデルは、微調整された場合、言語的特徴を置き換えるのに十分強力であることがわかった。
言語的な特徴を使わずに、モデルが知識集約型モデルより優れているのは、これが初めてである。
論文 参考訳(メタデータ) (2020-06-21T17:18:01Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。