8 papers · 1 filter
Exploring Extrinsic and Intrinsic Properties for Effective Reasoning with Code Interpreter
Patomporn Payoungkhamdee, Napat Laosaengpha, Jenta Wonglertsakul +8
Reasoning with a Code Interpreter (CI) has emerged as an effective paradigm for enhancing the reasoning capabilities of large language models (LLMs) through executable computation…
SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?
Wuttikorn Ponwitayarat, Peerat Limkonchotiwat, Raymond Ng +9
Multilingual text embeddings are often assumed to encode meaning in a perspective-independent semantic space, yielding stable similarity judgments across tasks and languages. Our r…
SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages
Muhammad Dehan Al Kautsar, Aswin Candra, Muhammad Alif Al Hakim +6
Although numerous datasets have been developed to support dialogue systems, most existing chit-chat datasets overlook the cultural nuances inherent in natural human conversations.…
Mangosteen: An Open Thai Corpus for Language Model Pretraining
Wannaphong Phatthiyaphaibun, Can Udomcharoenchaikit, Pakpoom Singkorapoom +4
Pre-training data shapes a language model's quality, but raw web text is noisy and demands careful cleaning. Existing large-scale corpora rely on English-centric or language-agnost…
Can Group Relative Policy Optimization Improve Thai Legal Reasoning and Question Answering?
Pawitsapak Akarajaradwong, Chompakorn Chaksangchaichot, Pirat Pothavorn +3
The Retrieval-Augmented Generation (RAG) systems' performance on Thai legal question answering is still limited, especially for questions requiring extensive, complex legal reasoni…
Explainable Depression Detection using Masked Hard Instance Mining
Patawee Prakrankamanant, Shinji Watanabe, Ekapol Chuangsuwanich
This paper addresses the critical need for improved explainability in text-based depression detection. While offering predictive outcomes, current solutions often overlook the unde…