10 papers · 1 filter
SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages
Muhammad Dehan Al Kautsar, Aswin Candra, Muhammad Alif Al Hakim +6
Although numerous datasets have been developed to support dialogue systems, most existing chit-chat datasets overlook the cultural nuances inherent in natural human conversations.…
SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?
Wuttikorn Ponwitayarat, Peerat Limkonchotiwat, Raymond Ng +9
Multilingual text embeddings are often assumed to encode meaning in a perspective-independent semantic space, yielding stable similarity judgments across tasks and languages. Our r…
Spatial Language Likelihood Grounding Network for Bayesian Fusion of Human-Robot Observations
Supawich Sitdhipol, Waritwong Sukprasongdee, Ekapol Chuangsuwanich +1
Fusing information from human observations can help robots overcome sensing limitations in collaborative tasks. However, an uncertainty-aware fusion framework requires a grounded l…
Mangosteen: An Open Thai Corpus for Language Model Pretraining
Wannaphong Phatthiyaphaibun, Can Udomcharoenchaikit, Pakpoom Singkorapoom +4
Pre-training data shapes a language model's quality, but raw web text is noisy and demands careful cleaning. Existing large-scale corpora rely on English-centric or language-agnost…
Can Group Relative Policy Optimization Improve Thai Legal Reasoning and Question Answering?
Pawitsapak Akarajaradwong, Chompakorn Chaksangchaichot, Pirat Pothavorn +3
The Retrieval-Augmented Generation (RAG) systems' performance on Thai legal question answering is still limited, especially for questions requiring extensive, complex legal reasoni…
THAI Speech Emotion Recognition (THAI-SER) corpus
Jilamika Wongpithayadisai, Chompakorn Chaksangchaichot, Soravitt Sangnark +7
We present the first sizeable corpus of Thai speech emotion recognition, THAI-SER, containing 41 hours and 36 minutes (27,854 utterances) from 100 recordings made in different reco…