CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia
Abstract Overview
CultureConverse is a multilingual simulation and evaluation harness for culturally grounded assistant dialogue across East and Southeast Asia. Rather than testing single-turn factual recall via multiple-choice questions, it evaluates multi-turn help-seeking interactions where the assistant must infer hidden cultural and taboo constraints from partial user disclosure. The framework covers 10 regions, 58 subgroup identities, and 7 domains, structuring simulation and scoring through disjoint public, private, and oracle context views. The released CultureConverse-DS dataset includes 14,610 benchmark evaluation episodes and 274,295 oracle-guided training dialogues, evaluating models on assistance quality (Helpfulness, Honesty, Harmlessness) and dialogue realism.
Novelty
The framework shifts cultural evaluation from static factual recall to dynamic, multi-turn assistant interactions governed by progressive disclosure and hidden cultural constraints. It integrates broad regional and intersectional subgroup coverage in East and Southeast Asia with an auditable, multi-stage generation pipeline grounded in cultural and taboo knowledge bases.
Results
Across 18 evaluated models, the benchmark effectively measures culturally grounded assistance, with GPT-5 mini achieving the highest average 3H score (4.28) and GPT-5.4 reaching the highest clean rate (91.0%). Human validation across 42 region-matched annotators showed that the automated judge achieved 90.1% agreement and a 0.581 MAE against human consensus. Fine-tuning two 8B models on 27,860 high-quality CultureConverse-DS episodes improved in-domain assistance and produced modest out-of-domain transfer gains on external cultural MCQ and safety classification benchmarks.
Key Points
- CultureConverse evaluates culturally grounded assistance through multi-turn dialogue under partial observability and progressive disclosure rather than static factual recall.
- The benchmark covers 10 East and Southeast Asian regions, 58 subgroup identities, and 7 domains, releasing 14,610 evaluation episodes alongside 274,295 oracle-guided training dialogues.
- Experiments demonstrate strong automated judge alignment with human consensus and show that fine-tuning on high-quality episodes yields modest gains across both in-domain and external cultural benchmarks.