2 papers
cs.CV2026
ITSELF: Attention Guided Fine-Grained Alignment for Vision-Language Retrieval
Tien-Huy Nguyen, Huu-Loc Tran, Thanh Duc Ngo
Vision Language Models (VLMs) have rapidly advanced and show strong promise for text-based person search (TBPS), a task that requires capturing fine-grained relationships between i…
cs.CV2024
Stratified Domain Adaptation: A Progressive Self-Training Approach for Scene Text Recognition
Kha Nhat Le, Hoang-Tuan Nguyen, Hung Tien Tran +1
Unsupervised domain adaptation (UDA) has become increasingly prevalent in scene text recognition (STR), especially where training and testing data reside in different domains. The…