Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
MulCLIP: A Multi-level Alignment Framework for Enhancing Fine-grained Long-context CLIP
Chau Truong, Hieu Ta Quang, Dung D. Le
Vision-language models like CLIP show impressive ability to align images and text, but their training on short, concise captions makes them struggle with lengthy, detailed descript…
cs.CV2025
Robustness Evaluation of OCR-based Visual Document Understanding under Multi-Modal Adversarial Attacks
Dong Nguyen Tien, Dung D. Le
Visual Document Understanding (VDU) systems have achieved strong performance in information extraction by integrating textual, layout, and visual signals. However, their robustness…