activity
20242026
collaborators

5 papers

cs.CV2026

Image Captioning via Compact Bidirectional Architecture

Zijie Song, Yuanen Zhou, Zhenzhen Hu +4

Most current image captioning models typically generate captions from left-to-right. This unidirectional property makes them can only leverage past context but not future context.…

cs.CV2025

Static for Dynamic: Towards a Deeper Understanding of Dynamic Facial Expressions Using Static Expression Data

Yin Chen, Jia Li, Yu Zhang +4

Dynamic facial expression recognition (DFER) infers emotions from the temporal evolution of expressions, unlike static facial expression recognition (SFER), which relies solely on…

cs.CV2025

Seeing is Believing? Enhancing Vision-Language Navigation using Visual Perturbations

Xuesong Zhang, Jia Li, Yunbo Xu +2

Autonomous navigation guided by natural language instructions in embodied environments remains a challenge for vision-language navigation (VLN) agents. Although recent advancements…

cs.CV2025

Grid Jigsaw Representation with CLIP: A New Perspective on Image Clustering

Zijie Song, Zhenzhen Hu, Richang Hong

Unsupervised representation learning for image clustering is essential in computer vision. Although the advancement of visual models has improved image clustering with efficient vi…

cs.CV2024

Text Proxy: Decomposing Retrieval from a 1-to-N Relationship into N 1-to-1 Relationships for Text-Video Retrieval

Jian Xiao, Zhenzhen Hu, Jia Li +1

Text-video retrieval (TVR) has seen substantial advancements in recent years, fueled by the utilization of pre-trained models and large language models (LLMs). Despite these advanc…