activity
20242026
collaborators
Showing cs.CLShow all

18 papers · 1 filter

cs.CL2026

A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays

Pawitsapak Akarajaradwong, Wuttikrai Lertprasertphakorn, Chompakorn Chaksangchaichot +1

Free-form legal essay evaluation in NLP treats expert inter-rater stability as a single ceiling number, and treats LLM-judge agreement with that ceiling as evidence of judge stabil…

cs.CL2026

Exploring Extrinsic and Intrinsic Properties for Effective Reasoning with Code Interpreter

Patomporn Payoungkhamdee, Napat Laosaengpha, Jenta Wonglertsakul +8

Reasoning with a Code Interpreter (CI) has emerged as an effective paradigm for enhancing the reasoning capabilities of large language models (LLMs) through executable computation…

cs.CL2026

DuDi: Dual-Signal Distillation with Cross-Lingual Verbalizer

Patomporn Payoungkhamdee, Tinnakit Udsa, Jian Gang Ngui +3

Small language models (SLMs) are efficient and scalable, but their multilingual capabilities degrade severely at sub-billion scales, especially for Southeast Asian (SEA) languages.…

cs.CL2026

SEA-NLI: Natural Language Inference as a Lens into Southeast Asian Cultural Understanding

Peerawat Chomphooyod, Jian Gang Ngui, Yosephine Susanto +5

Frontier LLMs perform well in Western contexts, but remain poorly tested on underrepresented cultures such as those in Southeast Asia (SEA). Existing NLI benchmarks are largely Wes…

cs.CL2026

SEA-Embedding: Open and Reproducible Text Embeddings for Southeast Asia

Peerat Limkonchotiwat, Raymond Ng, Sarana Nutanong +1

Text embeddings are fundamental to many downstream applications, making robustness important for real-world NLP. However, most recent state-of-the-art embedding models are not repr…

cs.CL2026

SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?

Wuttikorn Ponwitayarat, Peerat Limkonchotiwat, Raymond Ng +9

Multilingual text embeddings are often assumed to encode meaning in a perspective-independent semantic space, yielding stable similarity judgments across tasks and languages. Our r…