most citedGlobal MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

6 citations · 6 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CL20246 cited

Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

Shivalika Singh, Angelika Romanou, Clémentine Fourrier +21

Cultural biases in multilingual datasets pose significant challenges for their effectiveness as global benchmarks. These biases stem not only from differences in language but also…

cs.CL2024

Kalahi: A handcrafted, grassroots cultural LLM evaluation suite for Filipino

Jann Railey Montalan, Jian Gang Ngui, Wei Qi Leong +4

Multilingual large language models (LLMs) today may not necessarily provide culturally appropriate and relevant responses to its Filipino users. We introduce Kalahi, a cultural LLM…

cs.CL2024

ThaiCoref: Thai Coreference Resolution Dataset

Pontakorn Trakuekul, Wei Qi Leong, Charin Polpanumas +3

While coreference resolution is a well-established research area in Natural Language Processing (NLP), research focusing on Thai language remains limited due to the lack of large a…

cs.CL2024

SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages

Holy Lovenia, Rahmad Mahendra, Salsabil Maulana Akbar +58

Southeast Asia (SEA) is a region rich in linguistic diversity and cultural variety, with over 1,300 indigenous languages and a population of 671 million people. However, prevailing…

cs.CL2024

Thai Universal Dependency Treebank

Panyut Sriwirote, Wei Qi Leong, Charin Polpanumas +4

Automatic dependency parsing of Thai sentences has been underexplored, as evidenced by the lack of large Thai dependency treebanks with complete dependency structures and the lack…