activity
20242026
collaborators

14 papers

cs.CL2026

ViTOED: A Dataset for Target-Oriented Emotion Detection on Vietnamese Social Media Texts

Chanh Vo, Son T. Luu, Ngan Luu-Thuy Nguyen

This paper introduces ViTOED, a novel dataset for target-oriented emotion detection in Vietnamese social media texts. The ViTOED comprises 10,985 user comments and 21,244 manually…

cs.CL2026

Syllabic-Structure Decoder for Automatic Speech Recognition in Vietnamese

Nghia Hieu Nguyen, Quan Ngoc Hoang, Long Hoang Huu Nguyen +2

Most Automatic Speech Recognition (ASR) systems formulate transcription as a prediction problem over orthographic units such as characters, subwords, or words. Although effective,…

cs.CL2026

Phonetic Modeling of Dialectal Variation in Vietnamese Speech

Quan Ngoc Hoang, Long Hoang Huu Nguyen, Nghia Hieu Nguyen +2

Vietnamese exhibits substantial dialectal phonetic variation across Northern, Central, and Southern regions, where identical lexical items may be realized with markedly different p…

cs.CV2026

Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention

Nhi Ngoc-Yen Nguyen, Anh-Duc Nguyen, Nghia Hieu Nguyen +2

Scene-text image captioning requires fusing three information streams -- visual features, OCR-detected text, and linguistic knowledge -- to generate descriptions that faithfully in…

cs.CL2026

ViCLSR: A Supervised Contrastive Learning Framework with Natural Language Inference for Natural Language Understanding Tasks

Tin Van Huynh, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen

High-quality text representations are crucial for natural language understanding (NLU), but low-resource languages like Vietnamese face challenges due to limited annotated data. Wh…

cs.CL2026

ViTextVQA: A Large-Scale Visual Question Answering Dataset and a Novel Multimodal Feature Fusion Method for Vietnamese Text Comprehension in Images

Quan Van Nguyen, Dan Quang Tran, Huy Quang Pham +4

Visual Question Answering (VQA) is a challenging task that requires the joint understanding of natural language and visual content. While early research primarily focused on recogn…