論文の概要: Beyond Cultural Knowledge: Evaluating Arabic Cultural Appropriateness of Large Language Models
- arxiv url: http://arxiv.org/abs/2609.16006v1
- Date: Thu, 30 Jul 2026 07:50:10 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-28 05:09:55.532757
- Title: Beyond Cultural Knowledge: Evaluating Arabic Cultural Appropriateness of Large Language Models
- Title(参考訳): 文化的知識を超えて:大規模言語モデルのアラビア文化的適切性を評価する
- Abstract要約: AraBehave: 1,623の文化的根拠とオープンエンドアラビア語のプロンプト,29,214の文化的適切性判定。
文化的適切性は、単一の能力ではなく、主に独立した2つの構成要素(規範的スタンスと基礎的文化的正確性)に分解される。
- 参考スコア(独自算出の注目度): 7.632624212809001
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Large language models (LLMs) increasingly serve users whose expectations are shaped by their cultural context, yet most cultural evaluations test what a model knows rather than how it behaves when giving open-ended recommendations, opinions, and guidance. We introduce AraBehave: 1,623 culturally grounded, open-ended Arabic prompts with 29,214 cultural-appropriateness judgments from native speakers across several Arab regions, plus a scoring model whose predictions correlate strongly with human judgments on unseen systems (Pearson r=0.74). Evaluating three Arabic-centric and three frontier LLMs, we find that cultural appropriateness is not a single capability but decomposes into two largely independent components: normative stance and grounded cultural accuracy. The best general-purpose and best Arabic-centric models score identically (3.84 vs. 3.83 of 5) yet almost never fail for the same reason: general-purpose models exhibit strong factual grounding but a culturally inappropriate normative stance, being penalized for secular framing and false balance on culturally settled matters (28--33% of their low-score rationales), while the best Arabic-centric model adopts the expected stance but is penalized for fabricated hadith and misquoted verses (29%). Stance is cheap and fragile: one sentence of cultural instruction lifts Gemini to 4.57, above every Arabic-specialized model. Conversely, a generic ``answer clearly and objectively'' prompt costs Allam-7B 0.68 points, while asking the same questions in English lowers scores for every model but one. Grounding instead tracks scale and Arabic alignment data, and disappears when culturally aware instruction tuning is replaced by a culture-neutral corpus. General safety benchmarks see none of this: they saturate above 89 while cultural scores span 2.71-3.84. We will release the benchmark, annotations, and the scoring model.
- Abstract(参考訳): 大きな言語モデル(LLM)は、期待が文化的文脈によって形作られたユーザに対して、ますます役立っているが、ほとんどの文化的評価は、オープンな推奨、意見、ガイダンスを与える際に、モデルがどのように振る舞うかではなく、モデルが知っていることをテストしている。
AraBehave: 1,623 の文化的根拠とオープンエンドアラビア語のプロンプト,29,214 のアラブ各地の先住民話者による文化的適切性判断に加えて,予測が見えないシステムの人的判断と強く相関するスコアモデルを紹介する(Pearson r=0.74)。
3つのアラビア中心のLLMと3つのフロンティアのLLMを評価すると、文化的適切性は単一の能力ではなく、規範的スタンスと文化的正確性という2つの大きな独立した構成要素に分解されることがわかった。
汎用モデルは、強い事実的根拠を示すが、文化的に不適切な規範的スタンスを示し、文化的に落ち着いた事柄に対する世俗的なフレーミングと虚偽のバランス(28~33%の低調な合理性)に罰せられる。
1つの文化的指導文はジェミニ語を4.57まで引き上げ、アラビア語に特化しているすべてのモデルより上である。
逆に、ジェネリックな ``Awer clearly and objectively'' プロンプトは Allam-7B 0.68 のコストがかかる。
グラウンドリングは、スケールとアラビアのアライメントデータを追跡し、文化的に認識された指導チューニングがカルチャーニュートラルコーパスに置き換えられると消滅する。
一般的な安全ベンチマークでは、89以上の飽和度を示し、文化的なスコアは2.71-3.84である。
ベンチマーク、アノテーション、スコアリングモデルをリリースします。
関連論文リスト
- Whose Norms? Disentangling Cultural and Personal Alignment in Large Language Models [66.34115263472931]
大規模言語モデルは、文化的規範と個人の嗜好のバランスを必要とする社会的意思決定の状況にますます使われてきている。
PACT(Personal-Preference and Cultural-Norm Trade-off framework)を導入し,モデルが文化的規範に従うか,あるいは個人の嗜好を許容するかを評価する。
PACTに関する5つの国による研究は、人間の文化のフォローは主にシナリオ国によって行われており、参加者が自身の文化的文脈を判断する際の合意は最低であることを示している。
論文 参考訳(メタデータ) (2026-06-05T22:20:30Z) - Commonsense Reasoning in Arab Culture [6.116784716369165]
我々は,現代標準アラビア語(MSA)における常識推論データセットであるArabCultureを紹介し,湾岸,レバント,北アフリカ,ナイルバレーの13カ国の文化をカバーしている。
ArabCultureは12の日常生活ドメインと54の細かいサブトピックにまたがっており、社会規範、伝統、日常生活の様々な側面を反映している。
ゼロショット評価は、最大32Bパラメータを持つオープンウェイト言語モデルは、様々なアラブ文化を理解するのに苦労し、地域によってパフォーマンスが異なることを示している。
論文 参考訳(メタデータ) (2025-02-18T11:49:54Z) - Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in
Large Language Models [89.94270049334479]
本稿では,大規模言語モデル(LLM)における文化的優位性について述べる。
LLMは、ユーザーが非英語で尋ねるときに期待する文化とは無関係な、不適切な英語文化関連の回答を提供することが多い。
論文 参考訳(メタデータ) (2023-10-19T05:38:23Z) - AceGPT, Localizing Large Language Models in Arabic [73.39989503874634]
本稿では,アラビア語のテキストによる事前学習,ネイティブなアラビア語命令を利用したSFT(Supervised Fine-Tuning),アラビア語のGPT-4応答を含む総合的なソリューションを提案する。
目標は、文化的に認知され、価値に整合したアラビア語のLLMを、多様で応用特有のアラビア語コミュニティのニーズに適応させることである。
論文 参考訳(メタデータ) (2023-09-21T13:20:13Z) - Having Beer after Prayer? Measuring Cultural Bias in Large Language Models [25.722262209465846]
多言語およびアラビア語のモノリンガルLMは、西洋文化に関連する実体に対して偏見を示すことを示す。
アラブ文化と西洋文化を対比する8つのタイプにまたがる628個の自然発生プロンプトと20,368個のエンティティからなる新しい資源であるCAMeLを紹介した。
CAMeLを用いて、物語生成、NER、感情分析などのタスクにおいて、16の異なるLMのアラビア語における異文化間性能について検討した。
論文 参考訳(メタデータ) (2023-05-23T18:27:51Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。