Culture Cartography: Mapping the Landscape of Cultural Knowledge
By: Caleb Ziems , William Held , Jane Yu and more
Potential Business Impact:
Teaches computers about different cultures better.
To serve global users safely and productively, LLMs need culture-specific knowledge that might not be learned during pre-training. How do we find such knowledge that is (1) salient to in-group users, but (2) unknown to LLMs? The most common solutions are single-initiative: either researchers define challenging questions that users passively answer (traditional annotation), or users actively produce data that researchers structure as benchmarks (knowledge extraction). The process would benefit from mixed-initiative collaboration, where users guide the process to meaningfully reflect their cultures, and LLMs steer the process towards more challenging questions that meet the researcher's goals. We propose a mixed-initiative methodology called CultureCartography. Here, an LLM initializes annotation with questions for which it has low-confidence answers, making explicit both its prior knowledge and the gaps therein. This allows a human respondent to fill these gaps and steer the model towards salient topics through direct edits. We implement this methodology as a tool called CultureExplorer. Compared to a baseline where humans answer LLM-proposed questions, we find that CultureExplorer more effectively produces knowledge that leading models like DeepSeek R1 and GPT-4o are missing, even with web search. Fine-tuning on this data boosts the accuracy of Llama-3.1-8B by up to 19.2% on related culture benchmarks.
Similar Papers
Localized Cultural Knowledge is Conserved and Controllable in Large Language Models
Computation and Language
Makes computers speak other languages like locals.
Grounding Multilingual Multimodal LLMs With Cultural Knowledge
Computation and Language
Helps computers understand different cultures worldwide.
Grounding Multilingual Multimodal LLMs With Cultural Knowledge
Computation and Language
Helps AI understand different cultures and languages better.