activity
20182026
collaborators
Showing 2025Show all

10 papers · 1 filter

cs.CL2025

SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages

Muhammad Dehan Al Kautsar, Aswin Candra, Muhammad Alif Al Hakim +6

Although numerous datasets have been developed to support dialogue systems, most existing chit-chat datasets overlook the cultural nuances inherent in natural human conversations.…

cs.CL2025

SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?

Wuttikorn Ponwitayarat, Peerat Limkonchotiwat, Raymond Ng +9

Multilingual text embeddings are often assumed to encode meaning in a perspective-independent semantic space, yielding stable similarity judgments across tasks and languages. Our r…

cs.RO2025

Spatial Language Likelihood Grounding Network for Bayesian Fusion of Human-Robot Observations

Supawich Sitdhipol, Waritwong Sukprasongdee, Ekapol Chuangsuwanich +1

Fusing information from human observations can help robots overcome sensing limitations in collaborative tasks. However, an uncertainty-aware fusion framework requires a grounded l…

cs.CL2025

Mangosteen: An Open Thai Corpus for Language Model Pretraining

Wannaphong Phatthiyaphaibun, Can Udomcharoenchaikit, Pakpoom Singkorapoom +4

Pre-training data shapes a language model's quality, but raw web text is noisy and demands careful cleaning. Existing large-scale corpora rely on English-centric or language-agnost…

cs.CL2025

Can Group Relative Policy Optimization Improve Thai Legal Reasoning and Question Answering?

Pawitsapak Akarajaradwong, Chompakorn Chaksangchaichot, Pirat Pothavorn +3

The Retrieval-Augmented Generation (RAG) systems' performance on Thai legal question answering is still limited, especially for questions requiring extensive, complex legal reasoni…

cs.SD2025

THAI Speech Emotion Recognition (THAI-SER) corpus

Jilamika Wongpithayadisai, Chompakorn Chaksangchaichot, Soravitt Sangnark +7

We present the first sizeable corpus of Thai speech emotion recognition, THAI-SER, containing 41 hours and 36 minutes (27,854 utterances) from 100 recordings made in different reco…