3 papers
cs.CL2026
Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean
Phannet Pov, Sovandara Chhoun, Hyun Woo Park +2
Large pretrained text-to-speech (TTS) models sound almost human for well-resourced languages, but much worse for languages that are rare in their training data. We study this quali…
cs.CL2026
Evaluation of Chunking Strategies for Effective Text Embedding in Low-Resource Language on Agricultural Documents
Sovandara Chhoun, Pichdara Po, Sereiwathna Ros +2
In this study, we compare the performance of four text chunking approaches: Recursive, Khmer-Aware, Sentence-Based, and LLM-Based within a Retrieval-Augmented Generation (RAG) fram…
cs.CL2026
A Comparative Study of Language Models for Khmer Retrieval-Augmented Question Answering
Sereiwathna Ros, Phannet Pov, Ratanaktepi Chhor +3
Retrieval-Augmented Generation (RAG) has emerged as a promising paradigm for grounding large language model (LLM) outputs in retrieved evidence, thereby reducing hallucination and…