5 papers
SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?
Wuttikorn Ponwitayarat, Peerat Limkonchotiwat, Raymond Ng +9
Multilingual text embeddings are often assumed to encode meaning in a perspective-independent semantic space, yielding stable similarity judgments across tasks and languages. Our r…
SEA-Guard: Culturally Grounded Multilingual Safeguard for Southeast Asia
Panuthep Tasawong, Jian Gang Ngui, Alham Fikri Aji +2
Culturally aware safeguards are crucial for AI alignment in real-world settings, where safety extends beyond common sense and encompasses diverse local values, norms, and region-sp…
SEA-SafeguardBench: Evaluating AI Safety in SEA Languages and Cultures
Panuthep Tasawong, Jian Gang Ngui, Alham Fikri Aji +2
Safeguard models help large language models (LLMs) detect and block harmful content, but most evaluations remain English-centric and overlook linguistic and cultural diversity. Exi…
WangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thai
Peerat Limkonchotiwat, Pume Tuchinda, Lalita Lowphansirikul +5
Large language models excel at instruction-following in English, but their performance in low-resource languages like Thai remains underexplored. Existing benchmarks often rely on…
NitiBench: A Comprehensive Study of LLM Framework Capabilities for Thai Legal Question Answering
Pawitsapak Akarajaradwong, Pirat Pothavorn, Chompakorn Chaksangchaichot +4
The application of large language models (LLMs) in the legal domain holds significant potential for information retrieval and question answering, yet Thai legal QA systems face cha…