activity
20232026
most citedGlobal MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

6 citations · 10 across the 11 of their papers we have counts for

collaborators

11 papers

cs.CL2026

DuDi: Dual-Signal Distillation with Cross-Lingual Verbalizer

Patomporn Payoungkhamdee, Tinnakit Udsa, Jian Gang Ngui +3

Small language models (SLMs) are efficient and scalable, but their multilingual capabilities degrade severely at sub-billion scales, especially for Southeast Asian (SEA) languages.…

cs.CL2026

SEA-NLI: Natural Language Inference as a Lens into Southeast Asian Cultural Understanding

Peerawat Chomphooyod, Jian Gang Ngui, Yosephine Susanto +5

Frontier LLMs perform well in Western contexts, but remain poorly tested on underrepresented cultures such as those in Southeast Asia (SEA). Existing NLI benchmarks are largely Wes…

cs.CL2026

SEA-LION-Embedding: Open and Reproducible Text Embeddings for Southeast Asia

Peerat Limkonchotiwat, Raymond Ng, Sarana Nutanong +1

Text embeddings are fundamental to many downstream applications, making robustness important for real-world NLP. However, most recent state-of-the-art embedding models are not repr…

cs.CL2026

SEA-Guard: Culturally Grounded Multilingual Safeguard for Southeast Asia

Panuthep Tasawong, Jian Gang Ngui, Alham Fikri Aji +2

Culturally aware safeguards are crucial for AI alignment in real-world settings, where safety extends beyond common sense and encompasses diverse local values, norms, and region-sp…

cs.CL2026

BURMESE-SAN: Burmese NLP Benchmark for Evaluating Large Language Models

Thura Aung, Jann Railey Montalan, Jian Gang Ngui +1

We introduce BURMESE-SAN, the first holistic benchmark that systematically evaluates large language models (LLMs) for Burmese across three core NLP competencies: understanding (NLU…

cs.CL2025

SEA-SafeguardBench: Evaluating AI Safety in SEA Languages and Cultures

Panuthep Tasawong, Jian Gang Ngui, Alham Fikri Aji +2

Safeguard models help large language models (LLMs) detect and block harmful content, but most evaluations remain English-centric and overlook linguistic and cultural diversity. Exi…