5 papers
Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering
Hangrui Xu, Zhengxian Wu, Yunyao Yu +6
Knowledge-Based Visual Question Answering (KB-VQA) relies on retrieving external information to answer queries involving long-tail entities. However, existing retrieval pipelines p…
Learning to Search: A Decision-Based Agent for Knowledge-Based Visual Question Answering
Zhuohong Chen, Zhenxian Wu, Yunyao Yu +6
Knowledge-based visual question answering (KB-VQA) requires vision-language models to understand images and use external knowledge, especially for rare entities and long-tail facts…
PSGait: Gait Recognition using Parsing Skeleton
Hangrui Xu, Zhengxian Wu, Chuanrui Zhang +4
Gait recognition has emerged as a robust biometric modality due to its non-intrusive nature. Conventional gait recognition methods mainly rely on silhouettes or skeletons. While ef…
Language-Guided and Motion-Aware Gait Representation for Generalizable Recognition
Zhengxian Wu, Chuanrui Zhang, Shenao Jiang +6
Gait recognition is emerging as a promising technology and an innovative field within computer vision, with a wide range of applications in remote human identification. However, ex…
DAGait: Generalized Skeleton-Guided Data Alignment for Gait Recognition
Zhengxian Wu, Chuanrui Zhang, Hangrui Xu +2
Gait recognition is emerging as a promising and innovative area within the field of computer vision, widely applied to remote person identification. Although existing gait recognit…