論文の概要: Challenges in annotations by humans and LLMs: A case study of evaluative language
- arxiv url: http://arxiv.org/abs/2607.28119v1
- Date: Thu, 30 Jul 2026 12:28:54 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-31 21:37:00.551714
- Title: Challenges in annotations by humans and LLMs: A case study of evaluative language
- Title(参考訳): 人間とLLMによるアノテーションの課題:評価言語を事例として
- Authors: Mirela Imamovic, Aenne Cecilia Kristine Knierim, Khushi Pitroda, Ekaterina Lapshinova-Koltunski,
- Abstract要約: 学習における言語学者、訓練された言語学者、および大規模言語モデル(LLM)によって生成されたアノテーションを比較した。
我々は, 評価理論とその態度サブシステムに注目し, 影響, 判断, 評価のカテゴリー(クラス)を含む。
訓練された言語学者が行ったアノテーションと比較すると,モデルが最善であるのに対して,訓練中の言語学者は高い合意点に達していない。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models (LLMs) to find out if they struggle with complex linguistic phenomena in a similar way. For this purpose, we analyse evaluative language in spoken popular science discourse, with the example of a corpus of English TED talk transcripts. We focus on the Appraisal theory and its Attitude subsystem, including the categories (classes) of Affect, Judgement, and Appreciation. In this context, Appraisal theory is an example of a highly subjective annotation task, making it a suitable example for the study of complex annotation challenges. First, we assess human annotations on a sentence level in specific scientific domains. Then, we develop three prompts and compare them for model performance for the automatic classification of Appraisal classes. We assess the performance of three LLMs using the best-performing prompt and finetune the model, reaching an F1-score of 0.77. We find that models perform best compared to annotations conducted by the trained linguist, while linguists in training do not reach high agreement scores. We conclude that LLMs can aid in complex annotation task resolution, opening new pathways for the complex theories annotated and analyzed in digital humanities studies.
- Abstract(参考訳): 本稿では、訓練における言語学者、訓練された言語学者、および大規模言語モデル(LLM)が生成するアノテーションの比較を行い、それらが同様の方法で複雑な言語現象に苦しむかどうかを調べる。
この目的のために、英文TEDトークテキストのコーパスを例として、話し言葉の一般的な科学談話における評価言語を分析した。
我々は, 評価理論とその態度サブシステムに注目し, 影響, 判断, 評価のカテゴリー(クラス)を含む。
この文脈において、評価理論は高度に主観的なアノテーションタスクの例であり、複雑なアノテーション課題の研究に適した例である。
まず、特定の科学的領域における文レベルでの人間のアノテーションを評価する。
そこで我々は3つのプロンプトを開発し,評価クラスの自動分類のためのモデル性能の比較を行った。
最良性能のプロンプトを用いて3つのLLMの性能評価を行い,F1スコア0.77に到達した。
訓練された言語学者が行ったアノテーションと比較すると,モデルが最善であるのに対して,訓練中の言語学者は高い合意点に達していない。
我々は、LCMが複雑なアノテーションタスクの解決に役立ち、デジタル人文科学研究において注釈付けされ解析された複雑な理論の新たな経路を開くことができると結論付けた。
関連論文リスト
- Towards Corpus-Grounded Agentic LLMs for Multilingual Grammatical Analysis [0.5545791216381869]
本稿では, エージェント型大規模言語モデル (LLM) を用いて, 注釈付きコーパスの体系的解析を効率化する方法について検討する。
本稿では,自然言語タスク解釈などの概念を統合したコーパスグラウンド文法解析のためのエージェントフレームワークを提案する。
We test the system on multilingual grammatical tasks by the World Atlas of Language Structures (WALS) (英語)
論文 参考訳(メタデータ) (2025-11-28T21:27:58Z) - Comparing LLM Text Annotation Skills: A Study on Human Rights Violations in Social Media Data [2.812898346527047]
本研究では,ロシア語とウクライナ語におけるソーシャルメディア投稿のゼロショットおよび少数ショットアノテーションに対する大規模言語モデル(LLM)の機能について検討した。
これらのモデルの有効性を評価するため、それらのアノテーションは、人間の二重注釈付きラベルのゴールドスタンダードセットと比較される。
この研究は、各モデルが示すエラーと不一致のユニークなパターンを探求し、その強み、制限、言語間適応性に関する洞察を提供する。
論文 参考訳(メタデータ) (2025-05-15T13:10:47Z) - A linguistically-motivated evaluation methodology for unraveling model's abilities in reading comprehension tasks [10.181408678232055]
モデルのサイズやアーキテクチャに関わらず,特定の例が常に低いスコアを得られるという直感に基づいて,理解タスクを読むための評価手法を提案する。
この複雑さを特徴付けるためのセマンティックフレームアノテーションを活用し、モデルの難易度を考慮に入れうる7つの複雑さ要因について検討する。
論文 参考訳(メタデータ) (2025-01-29T11:05:20Z) - Holmes: A Benchmark to Assess the Linguistic Competence of Language Models [59.627729608055006]
言語モデル(LM)の言語能力を評価するための新しいベンチマークであるHolmesを紹介する。
我々は、計算に基づく探索を用いて、異なる言語現象に関するLMの内部表現を調べる。
その結果,近年,他の認知能力からLMの言語能力を引き離す声が上がっている。
論文 参考訳(メタデータ) (2024-04-29T17:58:36Z) - Disco-Bench: A Discourse-Aware Evaluation Benchmark for Language
Modelling [70.23876429382969]
本研究では,多種多様なNLPタスクに対して,文内談話特性を評価できるベンチマークを提案する。
ディスコ・ベンチは文学領域における9つの文書レベルのテストセットから構成されており、豊富な談話現象を含んでいる。
また,言語分析のために,対象モデルが談話知識を学習するかどうかを検証できる診断テストスイートを設計する。
論文 参考訳(メタデータ) (2023-07-16T15:18:25Z) - Syntax and Semantics Meet in the "Middle": Probing the Syntax-Semantics
Interface of LMs Through Agentivity [68.8204255655161]
このような相互作用を探索するためのケーススタディとして,作用性のセマンティックな概念を提示する。
これは、LMが言語アノテーション、理論テスト、発見のためのより有用なツールとして役立つ可能性を示唆している。
論文 参考訳(メタデータ) (2023-05-29T16:24:01Z) - Using Natural Language Explanations to Rescale Human Judgments [81.66697572357477]
大規模言語モデル(LLM)を用いて順序付けアノテーションと説明を再スケールする手法を提案する。
我々は、アノテータのLikert評価とそれに対応する説明をLLMに入力し、スコア付けルーリックに固定された数値スコアを生成する。
提案手法は,合意に影響を及ぼさずに生の判断を再スケールし,そのスコアを同一のスコア付けルーリックに接する人間の判断に近づける。
論文 参考訳(メタデータ) (2023-05-24T06:19:14Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。