activity
20242026
collaborators

8 papers

cs.CV2026

Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention

Nhi Ngoc-Yen Nguyen, Anh-Duc Nguyen, Nghia Hieu Nguyen +2

Scene-text image captioning requires fusing three information streams -- visual features, OCR-detected text, and linguistic knowledge -- to generate descriptions that faithfully in…

cs.CL2026

ViGoEmotions: A Benchmark Dataset For Fine-grained Emotion Detection on Vietnamese Texts

Hung Quang Tran, Nam Tien Pham, Son T. Luu +1

Emotion classification plays a significant role in emotion prediction and harmful content detection. Recent advancements in NLP, particularly through large language models (LLMs),…

cs.CL2026

ViCLSR: A Supervised Contrastive Learning Framework with Natural Language Inference for Natural Language Understanding Tasks

Tin Van Huynh, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen

High-quality text representations are crucial for natural language understanding (NLU), but low-resource languages like Vietnamese face challenges due to limited annotated data. Wh…

cs.IR2025

Optimizing Legal Document Retrieval in Vietnamese with Semi-Hard Negative Mining

Van-Hoang Le, Duc-Vu Nguyen, Kiet Van Nguyen +1

Large Language Models (LLMs) face significant challenges in specialized domains like law, where precision and domain-specific knowledge are critical. This paper presents a streamli…

cs.CL2025

ViMRHP: A Vietnamese Benchmark Dataset for Multimodal Review Helpfulness Prediction via Human-AI Collaborative Annotation

Truc Mai-Thanh Nguyen, Dat Minh Nguyen, Son T. Luu +1

Multimodal Review Helpfulness Prediction (MRHP) is an essential task in recommender systems, particularly in E-commerce platforms. Determining the helpfulness of user-generated rev…

cs.CL2025

LiGT: Layout-infused Generative Transformer for Visual Question Answering on Vietnamese Receipts

Thanh-Phong Le, Trung Le Chi Phan, Nghia Hieu Nguyen +1

Document Visual Question Answering (Document VQA) challenges multimodal systems to holistically handle textual, layout, and visual modalities to provide appropriate answers. Docume…