5 papers
Image Captioning via Compact Bidirectional Architecture
Zijie Song, Yuanen Zhou, Zhenzhen Hu +4
Most current image captioning models typically generate captions from left-to-right. This unidirectional property makes them can only leverage past context but not future context.…
Static for Dynamic: Towards a Deeper Understanding of Dynamic Facial Expressions Using Static Expression Data
Yin Chen, Jia Li, Yu Zhang +4
Dynamic facial expression recognition (DFER) infers emotions from the temporal evolution of expressions, unlike static facial expression recognition (SFER), which relies solely on…
Seeing is Believing? Enhancing Vision-Language Navigation using Visual Perturbations
Xuesong Zhang, Jia Li, Yunbo Xu +2
Autonomous navigation guided by natural language instructions in embodied environments remains a challenge for vision-language navigation (VLN) agents. Although recent advancements…
Grid Jigsaw Representation with CLIP: A New Perspective on Image Clustering
Zijie Song, Zhenzhen Hu, Richang Hong
Unsupervised representation learning for image clustering is essential in computer vision. Although the advancement of visual models has improved image clustering with efficient vi…
Text Proxy: Decomposing Retrieval from a 1-to-N Relationship into N 1-to-1 Relationships for Text-Video Retrieval
Jian Xiao, Zhenzhen Hu, Jia Li +1
Text-video retrieval (TVR) has seen substantial advancements in recent years, fueled by the utilization of pre-trained models and large language models (LLMs). Despite these advanc…