6 citations · 10 across the 11 of their papers we have counts for
11 papers
DuDi: Dual-Signal Distillation with Cross-Lingual Verbalizer
Patomporn Payoungkhamdee, Tinnakit Udsa, Jian Gang Ngui +3
Small language models (SLMs) are efficient and scalable, but their multilingual capabilities degrade severely at sub-billion scales, especially for Southeast Asian (SEA) languages.…
SEA-NLI: Natural Language Inference as a Lens into Southeast Asian Cultural Understanding
Peerawat Chomphooyod, Jian Gang Ngui, Yosephine Susanto +5
Frontier LLMs perform well in Western contexts, but remain poorly tested on underrepresented cultures such as those in Southeast Asia (SEA). Existing NLI benchmarks are largely Wes…
SEA-LION-Embedding: Open and Reproducible Text Embeddings for Southeast Asia
Peerat Limkonchotiwat, Raymond Ng, Sarana Nutanong +1
Text embeddings are fundamental to many downstream applications, making robustness important for real-world NLP. However, most recent state-of-the-art embedding models are not repr…
SEA-Guard: Culturally Grounded Multilingual Safeguard for Southeast Asia
Panuthep Tasawong, Jian Gang Ngui, Alham Fikri Aji +2
Culturally aware safeguards are crucial for AI alignment in real-world settings, where safety extends beyond common sense and encompasses diverse local values, norms, and region-sp…
BURMESE-SAN: Burmese NLP Benchmark for Evaluating Large Language Models
Thura Aung, Jann Railey Montalan, Jian Gang Ngui +1
We introduce BURMESE-SAN, the first holistic benchmark that systematically evaluates large language models (LLMs) for Burmese across three core NLP competencies: understanding (NLU…
SEA-SafeguardBench: Evaluating AI Safety in SEA Languages and Cultures
Panuthep Tasawong, Jian Gang Ngui, Alham Fikri Aji +2
Safeguard models help large language models (LLMs) detect and block harmful content, but most evaluations remain English-centric and overlook linguistic and cultural diversity. Exi…