論文の概要: Automatic Thematic Indexing of Large Literary Corpora: A Machine Learning Approach to Voltaire's Complete Works
- arxiv url: http://arxiv.org/abs/2607.09316v1
- Date: Fri, 10 Jul 2026 11:54:02 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-13 16:48:13.432319
- Title: Automatic Thematic Indexing of Large Literary Corpora: A Machine Learning Approach to Voltaire's Complete Works
- Title(参考訳): 大文字コーパスの自動セマンティックインデクシング:ボルテールの完全作品に対する機械学習アプローチ
- Authors: Miguel Arana-Catania, Gillian Pink, Glenn Roe,
- Abstract要約: 主題の索引付けは、大規模な文学や歴史の版に学術的にアクセスするために不可欠である。
本稿では,機械学習のセマンティックインデクシングへの応用について検討する。
- 参考スコア(独自算出の注目度): 0.33985395340995606
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Thematic indexing -- the practice of assigning structured conceptual labels to sections of text -- is essential to scholarly access in large-scale literary and historical editions, yet it remains a largely manual, labour-intensive process. This paper explores the application of machine learning to automatic thematic indexing, using two substantial sub-corpora of the Complete Works of Voltaire as a test case: the Essai sur les mœurs et l'esprit des nations and the Questions sur l'Encyclopédie. The task is framed as a multi-label classification problem, in which a model must assign the set of index entries that a professional indexer would apply to a given page of text. We compare a range of approaches -- from encoder-based models with classification heads to generative large language models (LLMs) fine-tuned via Low-Rank Adaptation (LoRA) -- spanning model sizes from approximately 3 to 120 billion parameters. Our best-performing model, from the Mistral family in a 4-bit quantised configuration, achieves F1 scores of up to 0.67; we argue that these figures represent lower bounds, given the inherent subjectivity of professional indexing and the frequency with which model predictions prove semantically valid despite diverging from the print index. We further evaluate cross-corpus generalisation and conduct a detailed qualitative analysis of model behaviour on literary and rhetorical features of the source texts that prove particularly resistant to automated treatment. Our findings have implications for the broader challenge of providing structured thematic access to large-scale literary and historical corpora.
- Abstract(参考訳): セマンティックインデックス(Thematic indexing) - 構造化された概念ラベルをテキストのセクションに割り当てる慣行は、学術的に学術的に学術的アクセスに欠かせないものであるが、主に手作業、労働集約的なプロセスである。
本稿では,Voltaire 完全作品の2つの実質的なサブコーパス,すなわち Essai sur les m'urs et l'esprit des nation と Questions sur l'Encyclopédie を用いて,自動主題索引化への機械学習の適用について検討する。
このタスクはマルチラベルの分類問題であり、モデルが与えられたテキストページに対してプロのインデクサが適用すべきインデックスエントリのセットを割り当てなければならない。
我々は、エンコーダベースのモデルと分類ヘッドを含むモデルから、ローランド適応(LoRA)を介して微調整されたジェネレーティブな大規模言語モデル(LLM)まで、モデルのサイズをおよそ3から120億のパラメータで比較する。
4ビットの量子化構成でミストラル族から得られた最良の性能モデルは、F1スコアを最大0.67まで達成し、これらの数値はプロの索引付けの固有の主観性と、モデル予測が印刷インデックスから分岐したにもかかわらず意味論的に有効であることを示す頻度から、下界を表すと論じている。
さらに、クロスコーパスの一般化を評価し、特に自動処理に耐性を示す原文の文学的・修辞的特徴に対するモデル行動の詳細な定性的な分析を行う。
本研究は,大規模文芸・歴史コーパスに構造化的アクセスを提供するという,より広範な課題を示唆するものである。
関連論文リスト
- Automatic Reflection Level Classification in Hungarian Student Essays [0.3262230127283452]
ハンガリーの学生エッセイにおける自動反射レベル分類の総合的研究について紹介する。
我々は、複数の学年で収集された1,954の反射的エッセイからなる、専門家によるハンガリーの大規模なデータセットを使用している。
TF-IDFとセマンティック埋め込み機能を用いた古典的機械学習モデルと、文書レベルのリフレクション分類のために微調整されたハンガリー固有のトランスフォーマーモデルである。
論文 参考訳(メタデータ) (2026-05-04T09:44:50Z) - Beyond Holistic Scores: Automatic Trait-Based Quality Scoring of Argumentative Essays [15.895792302323883]
教育の文脈では、教師と学習者は解釈可能な特性レベルのフィードバックを必要とする。
本稿では,2つの相補的モデリングパラダイムを用いた特徴量に基づく自動弁別評価手法について検討する。
スコア・オーディナリティを明示的にモデル化することは、人間のレーダとの合意を著しく改善することを示します。
論文 参考訳(メタデータ) (2026-02-04T14:30:52Z) - Toward Purpose-oriented Topic Model Evaluation enabled by Large Language Models [0.8193467416247519]
トピック品質の4つの重要な側面にまたがる9つのLarge Language Models(LLM)ベースのメトリクスを利用する目的指向評価フレームワークを導入する。
このフレームワークは、敵対的およびサンプリングベースのプロトコルを通じて検証され、ニュース記事、学術出版物、ソーシャルメディア投稿にまたがるデータセットに適用される。
論文 参考訳(メタデータ) (2025-09-08T18:46:08Z) - Modelling and Classifying the Components of a Literature Review [0.0]
本稿では, 言語モデル(LLM)を用いて, ドメインの専門家が手動で注釈付けした700文と, 自動ラベル付けされた2,240文からなる新しいベンチマークを提案する。
この実験は、この挑戦的な領域における芸術の状態を前進させるいくつかの新しい洞察をもたらす。
論文 参考訳(メタデータ) (2025-08-06T11:30:07Z) - Are Large Language Models Good Classifiers? A Study on Edit Intent Classification in Scientific Document Revisions [62.12545440385489]
大規模言語モデル(LLM)は、テキスト生成の大幅な進歩をもたらしたが、分類タスクの強化の可能性はまだ未検討である。
生成と符号化の両方のアプローチを含む分類のための微調整LDMを徹底的に研究するためのフレームワークを提案する。
我々はこのフレームワークを編集意図分類(EIC)においてインスタンス化する。
論文 参考訳(メタデータ) (2024-10-02T20:48:28Z) - CiteFusion: An Ensemble Framework for Citation Intent Classification Harnessing Dual-Model Binary Couples and SHAP Analyses [1.6695303704829412]
CiteFusionは、SciCiteとACL-ARCという2つのベンチマークデータセット上のマルチクラスCitation Intent Classificationタスクに対処する。
実験の結果、CiteFusionは最先端のパフォーマンスを達成し、Macro-F1スコアはSciCiteで89.60%、ACL-ARCで76.24%であった。
我々は、SciCiteで開発されたCiteFusionモデルを利用して、引用意図を分類するWebベースのアプリケーションをリリースする。
論文 参考訳(メタデータ) (2024-07-18T09:29:33Z) - Large Language Models in the Workplace: A Case Study on Prompt
Engineering for Job Type Classification [58.720142291102135]
本研究では,実環境における職種分類の課題について検討する。
目標は、英語の求職が卒業生やエントリーレベルの地位に適切かどうかを判断することである。
論文 参考訳(メタデータ) (2023-03-13T14:09:53Z) - Towards Making the Most of Context in Neural Machine Translation [112.9845226123306]
我々は、これまでの研究がグローバルな文脈をはっきりと利用しなかったと論じている。
本研究では,各文の局所的文脈を意図的にモデル化する文書レベルNMTフレームワークを提案する。
論文 参考訳(メタデータ) (2020-02-19T03:30:00Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。