4 papers
MulCLIP: A Multi-level Alignment Framework for Enhancing Fine-grained Long-context CLIP
Chau Truong, Hieu Ta Quang, Dung D. Le
Vision-language models like CLIP show impressive ability to align images and text, but their training on short, concise captions makes them struggle with lengthy, detailed descript…
Multilingual LLM Prompting Strategies for Medical English-Vietnamese Machine Translation
Nhu Vo, Nu-Uyen-Phuong Le, Dung D. Le +2
Medical English-Vietnamese machine translation (En-Vi MT) is essential for healthcare access and communication in Vietnam, yet Vietnamese remains a low-resource and under-studied l…
Communication-Efficient and Accurate Approach for Aggregation in Federated Low-Rank Adaptation
Le-Tuan Nguyen, Minh-Duong Nguyen, Seon-Geun Jeong +2
With the rapid emergence of foundation models and the increasing need for fine-tuning across distributed environments, Federated Low-Rank Adaptation (FedLoRA) has recently gained s…
Robustness Evaluation of OCR-based Visual Document Understanding under Multi-Modal Adversarial Attacks
Dong Nguyen Tien, Dung D. Le
Visual Document Understanding (VDU) systems have achieved strong performance in information extraction by integrating textual, layout, and visual signals. However, their robustness…