LLMs trained primarily on Western data misalign with non-Western cultural values; this work provides a replicable framework for embedding country-specific values in low-resource languages through survey-grounded datasets and targeted fine-tuning.
This paper introduces LKValues, a resource suite for aligning large language models with Sri Lankan cultural values. The authors surveyed 205 Sri Lankans to identify 40 key societal values, created a 150k-instance instruction dataset in Sinhala and English, and built an evaluation benchmark.