A team of researchers has released LKValues, a comprehensive framework designed to teach artificial intelligence systems to respect Sri Lankan cultural norms and societal principles. The work highlights a critical gap in how major language models handle values and ethics across different cultures, particularly in low-resource linguistic communities.
According to arXiv, the research addresses a fundamental problem: existing large language models tend to embed Western perspectives when making decisions about what constitutes appropriate or ethical behavior. For multilingual societies like Sri Lanka, which operates across Sinhala, Tamil, and English, this creates serious misalignment between AI system outputs and local values. Current evaluation benchmarks and training approaches largely ignore these region-specific cultural contexts entirely.
Building a Foundation From Local Input
The researchers conducted a trilingual survey of 205 participants to identify which values matter most to Sri Lankan society. Rather than simply importing global ethical frameworks, they combined established international guidelines with locally derived principles generated by AI systems themselves. This hybrid approach produced 40 core societal values with majority endorsement from respondents.
Using these validated values, the team constructed two key resources. LKValuesIT provides a training dataset of 150,000 instruction examples derived from Sinhala and English news sources, each scenario designed to test how models handle culturally sensitive situations. LKValuesBench offers a separate evaluation suite with 1,000 test instances that measure whether models actually respect identified Sri Lankan values in practice.
Exposing Performance Gaps Across Models
When the researchers evaluated both proprietary systems and open-source language models using their new benchmark, results revealed consistent shortfalls. Even newer and larger models demonstrated notable weaknesses when handling low-resource language contexts and culturally specific values. The team then fine-tuned three open-weight base models: Qwen 3.5 in both 4 billion and 9 billion parameter versions, plus Aya-Expanse at 8 billion parameters.
Fine-tuning with LKValues data improved performance in both Sinhala and English across the Qwen family, reducing invalid outputs and narrowing gaps between languages. However, the improvements proved somewhat dependent on the model architecture itself, suggesting that different systems require tailored approaches to value alignment.
A Replicable Blueprint for Global AI Development
The significance of this work extends beyond Sri Lanka. The researchers designed their methodology as a replicable pipeline that other countries and cultural communities could adapt for their own contexts. As AI systems become increasingly deployed in non-Western markets, the ability to systematically embed local values into these models becomes essential for responsible deployment.
The team has made their dataset publicly available on GitHub, enabling other researchers and developers to build upon this foundation. This openness reflects growing recognition within the AI research community that addressing cultural bias requires collaborative, bottom-up approaches rather than top-down enforcement of universal principles. For regions with limited resources to develop their own AI systems independently, frameworks like LKValues offer a practical path toward more culturally responsive artificial intelligence.



