FuguReport

CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia

Authors Bryan Chen Zhengyu Tan, Weihua Zheng, Thong T. Doan, Bich Ngoc Doan, Jia Wang Peh, Xiaoyuan Yi, Jing Yao, Xing Xie, Nancy F. Chen, Zhengyuan Liu, JinYeong Bak, Wafi Shamdi, Soo Kai Chie, Liew Yu Siong, Aina Azyyati Binti Mohamad Rezal, Lew Yan Yan Vanessa, Huadan Wu, Dylan Raharja, Nadya Yuki Wangsajaya, Akane Fukushige, Kazushi Kato, Koji Inoue, Tatsuya Kawahara, Jaehyung Seo, Dongjun Kim, Seungyoon Lee, Zi Haur Pang, Rui Yang Tan, Charibeth Ko Cheng, Maria Regina Justina Estuar, Jann Railey Montalan, Pham Minh Duc, Roy Ka-Wei Lee
Affiliations Konkuk University / AI Singapore / A*STAR / Singapore University of Technology and Design / De La Salle University / Korea University / École Polytechnique Fédérale de Lausanne / China University of Petroleum / Microsoft / Upstage AI / Kyoto University / Ateneo de Manila University / Sungkyunkwan University / Nanyang Technological University / Universiti Brunei Darussalam
Categories Method / Dialogue Systems / Multilingual culturally grounded assistant simulation, Evaluation / Benchmarking / Cultural dialogue benchmark construction, Application / Multilingual NLP / East and Southeast Asia assistance
License CC BY 4.0

Abstract Overview

CultureConverse is a multilingual simulation and evaluation harness for culturally grounded assistant dialogue across East and Southeast Asia. Rather than testing single-turn factual recall via multiple-choice questions, it evaluates multi-turn help-seeking interactions where the assistant must infer hidden cultural and taboo constraints from partial user disclosure. The framework covers 10 regions, 58 subgroup identities, and 7 domains, structuring simulation and scoring through disjoint public, private, and oracle context views. The released CultureConverse-DS dataset includes 14,610 benchmark evaluation episodes and 274,295 oracle-guided training dialogues, evaluating models on assistance quality (Helpfulness, Honesty, Harmlessness) and dialogue realism.

Novelty

The framework shifts cultural evaluation from static factual recall to dynamic, multi-turn assistant interactions governed by progressive disclosure and hidden cultural constraints. It integrates broad regional and intersectional subgroup coverage in East and Southeast Asia with an auditable, multi-stage generation pipeline grounded in cultural and taboo knowledge bases.

Results

Across 18 evaluated models, the benchmark effectively measures culturally grounded assistance, with GPT-5 mini achieving the highest average 3H score (4.28) and GPT-5.4 reaching the highest clean rate (91.0%). Human validation across 42 region-matched annotators showed that the automated judge achieved 90.1% agreement and a 0.581 MAE against human consensus. Fine-tuning two 8B models on 27,860 high-quality CultureConverse-DS episodes improved in-domain assistance and produced modest out-of-domain transfer gains on external cultural MCQ and safety classification benchmarks.

Key Points

  1. CultureConverse evaluates culturally grounded assistance through multi-turn dialogue under partial observability and progressive disclosure rather than static factual recall.
  2. The benchmark covers 10 East and Southeast Asian regions, 58 subgroup identities, and 7 domains, releasing 14,610 evaluation episodes alongside 274,295 oracle-guided training dialogues.
  3. Experiments demonstrate strong automated judge alignment with human consensus and show that fine-tuning on high-quality episodes yields modest gains across both in-domain and external cultural benchmarks.

References

This page was created using generative AI such as GPT-5, Claude Opus 4, Gemini 3, Gemini 3.1 Flash Image, and their higher-end successor versions. No guarantee can be made regarding its contents.