論文の概要: Evaluating Large Language Models for Antisemitic Incident Classification
- arxiv url: http://arxiv.org/abs/2607.04890v1
- Date: Mon, 06 Jul 2026 10:14:53 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-07 22:26:30.109328
- Title: Evaluating Large Language Models for Antisemitic Incident Classification
- Title(参考訳): アンチセミティックインシデント分類のための大規模言語モデルの評価
- Authors: Karina Halevy, Julia Mendelsohn, Chan Young Park, Yulia Tsvetkov, Maarten Sap,
- Abstract要約: 本稿では, 有害事象検出の課題を紹介し, きめ細かいラベルによる反ユダヤ事象の報告の発見と分類を行うAIシステムの能力について検討する。
OpenAIのGPT-4oとMetaのLlama-3.2-3B-Instructing on datasets including antisemitic event descriptions from News article, civil society report, and official records。
- 参考スコア(独自算出の注目度): 64.21103964951531
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Addressing hate and violence in society requires timely detection of hateful events from public reporting, but automated identification of hateful events remains underexplored. We introduce the task of hateful event detection and investigate the ability of AI systems, specifically large language models (LLMs), to discover and classify reports of antisemitic events with fine-grained labels. We evaluate OpenAI's GPT-4o and Meta's Llama-3.2-3B-Instruct on multiple expert-annotated datasets containing antisemitic event descriptions from news articles, civil society reports, and official records. We show that LLMs, particularly GPT-4o, have potential for this task, but substantial improvement is needed. Providing clear term definitions and in-context examples in prompts can improve performance: definitions are most helpful for rhetoric-oriented events (e.g. classical antisemitic tropes), while examples help label action-oriented events (e.g. physical assault). A case study of college newspapers demonstrates that LLMs can help surface relevant real-world events, supporting early monitoring and intervention. Overall, our findings highlight both opportunities and critical gaps in AI's ability to recognize complex harms and underscore the need for collaborative efforts among AI developers, policymakers, and civil society to design models, implement robust evaluation, and develop policy frameworks for defining and combating hate efficiently and effectively.
- Abstract(参考訳): 社会における憎悪や暴力に対処するには、公的報告からヘイトフルな出来事をタイムリーに検出する必要があるが、ヘイトフルな出来事の自動識別はいまだに探索されていない。
本稿では,AIシステム,特に大規模言語モデル(LLM)の課題について紹介し,詳細なラベルによる反ユダヤ的事象の報告の発見と分類を行う。
OpenAIのGPT-4oとMetaのLlama-3.2-3B-Instructing on multiple expert-annotated datasets including antisemitic event descriptions from news article, civil society report, and official records。
LLM、特にGPT-4oは、この課題に可能性を持っているが、かなりの改善が必要である。
明確な用語の定義とインコンテキストの例を提供することで、パフォーマンスを向上させることができる: 定義はレトリック指向のイベント(例えば古典的な反セミティックな傾向)に最も役立ち、例はアクション指向のイベント(例えば物理的な攻撃)に役立ちます。
大学新聞のケーススタディでは、LDMが関連する現実世界の出来事を表面化し、早期の監視と介入を支援することを実証している。
全体として、私たちの発見は、複雑な害を認識し、AI開発者、政策立案者、市民社会がモデルの設計、堅牢な評価を実装し、ヘイトを効果的に定義し、戦うための政策フレームワークを開発することの必要性を、AIの能力の機会と重要なギャップの両方を強調しています。
関連論文リスト
- Impacts of Racial Bias in Historical Training Data for News AI [0.0]
本稿では,ニューヨーク・タイムズ・アノテート・コーパス(New York Times Annotated Corpus)で訓練された複数ラベルの分類器の作成について検討する。
ブラックス」ラベルは、一部のマイノリティ化されたグループにまたがる一般的な「人種差別検知器」として部分的に機能していることがわかった。
しかし、新型コロナウイルス時代の反アジア的ヘイトストーリーやブラック・ライブ・マター運動の報道など、現代の事例では期待に反する効果がある。
論文 参考訳(メタデータ) (2025-12-18T18:56:11Z) - Self-evolving expertise in complex non-verifiable subject domains: dialogue as implicit meta-RL [0.0]
いわゆる「邪悪な問題」は、複雑な多次元の設定、検証不可能な結果、不均一な影響、客観的に正しい答えの欠如など、歴史を通じて人類を悩ませてきた。
現状の人工知能システム(特にLarge Language Modelベースのエージェント)は、そのような問題を解決するために人間と共同で研究されている。
この研究は、Dialecticaとのギャップに対処する。これは、エージェントが定義されたトピックに関する構造化された対話に従事し、メモリによる拡張、自己回帰、ポリシーに制約のあるコンテキスト編集を行うフレームワークである。
論文 参考訳(メタデータ) (2025-10-17T15:59:44Z) - From Incidents to Insights: Patterns of Responsibility following AI Harms [1.9389881806157316]
AIインシデントデータベースは航空安全データベースにインスパイアされ、障害からの集合的学習を可能とし、将来のインシデントを防ぐ。
データベースは、ニュースやメディアから収集された数百のAI障害を文書化している。
技術的に焦点を絞った学習を超えて、データセットは新たな、非常に価値のある洞察を提供することができる、と私たちは主張する。
論文 参考訳(メタデータ) (2025-05-07T09:59:36Z) - Contextualize-then-Aggregate: Circuits for In-Context Learning in Gemma-2 2B [51.74607395697567]
In-Context Learning (ICL)は、大規模言語モデル(LLM)の興味深い能力である。
我々は5つの自然主義ICLタスクに対してGemma-2 2Bにおける情報フローを因果介入を用いて同定する。
このモデルでは,2段階戦略を用いてタスク情報を推論し,コンテキスト化-then-aggregateと呼ぶ。
論文 参考訳(メタデータ) (2025-03-31T18:33:55Z) - Investigating Annotator Bias in Large Language Models for Hate Speech Detection [5.589665886212444]
本稿では,ヘイトスピーチデータに注釈をつける際に,Large Language Models (LLMs) に存在するバイアスについて考察する。
具体的には、これらのカテゴリ内の非常に脆弱なグループを対象として、アノテータバイアスを分析します。
我々は,この研究を行うために,独自のヘイトスピーチ検出データセットであるHateBiasNetを紹介した。
論文 参考訳(メタデータ) (2024-06-17T00:18:31Z) - Survey of Vulnerabilities in Large Language Models Revealed by
Adversarial Attacks [5.860289498416911]
大規模言語モデル(LLM)はアーキテクチャと能力において急速に進歩しています。
複雑なシステムに深く統合されるにつれて、セキュリティ特性を精査する緊急性が高まっている。
本稿では,LSMに対する対人攻撃の新たな学際的分野について調査する。
論文 参考訳(メタデータ) (2023-10-16T21:37:24Z) - Bias and Fairness in Large Language Models: A Survey [73.87651986156006]
本稿では,大規模言語モデル(LLM)のバイアス評価と緩和手法に関する総合的な調査を行う。
まず、自然言語処理における社会的偏見と公平性の概念を統合し、形式化し、拡張する。
次に,3つの直感的な2つのバイアス評価法と1つの緩和法を提案し,文献を統一する。
論文 参考訳(メタデータ) (2023-09-02T00:32:55Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。