collaborators

7 papers

cs.CL2026

myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASR

Ye Kyaw Thu, Ye Bhone Lin, Thura Aung +6

Although Whisper models benefit from large-scale multilingual pre-training, their performance on Burmese medical speech remains limited. This work presents a Burmese medical speech…

cs.CL2026

BURMESE-SAN: Burmese NLP Benchmark for Evaluating Large Language Models

Thura Aung, Jann Railey Montalan, Jian Gang Ngui +1

We introduce BURMESE-SAN, the first holistic benchmark that systematically evaluates large language models (LLMs) for Burmese across three core NLP competencies: understanding (NLU…

cs.CL2026

SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?

Wuttikorn Ponwitayarat, Peerat Limkonchotiwat, Raymond Ng +9

Multilingual text embeddings are often assumed to encode meaning in a perspective-independent semantic space, yielding stable similarity judgments across tasks and languages. Our r…

cs.CL2025

ASR Error Correction in Low-Resource Burmese with Alignment-Enhanced Transformers using Phonetic Features

Ye Bhone Lin, Thura Aung, Ye Kyaw Thu +1

This paper investigates sequence-to-sequence Transformer models for automatic speech recognition (ASR) error correction in low-resource Burmese, focusing on different feature integ…

cs.CL2025

Enhancing Burmese News Classification with Kolmogorov-Arnold Network Head Fine-tuning

Thura Aung, Eaint Kay Khaing Kyaw, Ye Kyaw Thu +2

In low-resource languages like Burmese, classification tasks often fine-tune only the final classification layer, keeping pre-trained encoder weights frozen. While Multi-Layer Perc…

cs.CL2025

KAConvText: Novel Approach to Burmese Sentence Classification using Kolmogorov-Arnold Convolution

Ye Kyaw Thu, Thura Aung, Thazin Myint Oo +1

This paper presents the first application of Kolmogorov-Arnold Convolution for Text (KAConvText) in sentence classification, addressing three tasks: imbalanced binary hate speech d…