activity
20222026
most citedOpenViVQA: Task, Dataset, and Multimodal Fusion Models for Visual Question Answering in Vietnamese

35 citations · 62 across the 16 of their papers we have counts for

collaborators

16 papers

cs.CL2026

Beyond Atomic Tokens: Factorizing Syllables for Language Model Pretraining

Nghia Hieu Nguyen, Thai Bao Huynh, Binh-An Dinh-Le +4

Conventional tokenizers represent text as characters or statistically derived subwords, overlooking the internal phonological structure of syllables and often requiring large vocab…

cs.CL2026

Direct Image-to-Modern Vietnamese Translation of Han-Nom Manuscripts via Multimodal RLHF Preference Alignment

Thi Kim Trang Vo, Nghia Hieu Nguyen, Ha Minh Tan

Translating Han-Nom manuscripts into modern Vietnamese is challenging because historical pages are often degraded, the script contains rare logographic characters, and parallel sup…

cs.CL2026

Syllabic-Structure Decoder for Automatic Speech Recognition in Vietnamese

Nghia Hieu Nguyen, Quan Ngoc Hoang, Long Hoang Huu Nguyen +2

Most Automatic Speech Recognition (ASR) systems formulate transcription as a prediction problem over orthographic units such as characters, subwords, or words. Although effective,…

cs.CL2026

Phonetic Modeling of Dialectal Variation in Vietnamese Speech

Quan Ngoc Hoang, Long Hoang Huu Nguyen, Nghia Hieu Nguyen +2

Vietnamese exhibits substantial dialectal phonetic variation across Northern, Central, and Southern regions, where identical lexical items may be realized with markedly different p…

cs.CV2026

Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention

Nhi Ngoc-Yen Nguyen, Anh-Duc Nguyen, Nghia Hieu Nguyen +2

Scene-text image captioning requires fusing three information streams -- visual features, OCR-detected text, and linguistic knowledge -- to generate descriptions that faithfully in…

cs.CL2026

ViMultiChoice: Toward a Method That Gives Explanation for Multiple-Choice Reading Comprehension in Vietnamese

Trung Tien Cao, Lam Minh Thai, Nghia Hieu Nguyen +2

Multiple-choice Reading Comprehension (MCRC) models aim to select the correct answer from a set of candidate options for a given question. However, they typically lack the ability…