2 papers
cs.CL2025
SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?
Wuttikorn Ponwitayarat, Peerat Limkonchotiwat, Raymond Ng +9
Multilingual text embeddings are often assumed to encode meaning in a perspective-independent semantic space, yielding stable similarity judgments across tasks and languages. Our r…
cs.CL2025
Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation
Peerat Limkonchotiwat, Kanruethai Masuk, Surapon Nonesung +4
Large language models show promising results in various NLP tasks. Despite these successes, the robustness and consistency of LLMs in underrepresented languages remain largely unex…