7 papers
R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation
Zhuohong Chen, Zhengxian Wu, Zirui Liao +6
Vision-centric retrieval for VQA requires retrieving images to supply missing visual cues and integrating them into the reasoning process. However, selecting the right images and i…
Learning to Search: A Decision-Based Agent for Knowledge-Based Visual Question Answering
Zhuohong Chen, Zhenxian Wu, Yunyao Yu +6
Knowledge-based visual question answering (KB-VQA) requires vision-language models to understand images and use external knowledge, especially for rare entities and long-tail facts…
Stabilizing Unsupervised Self-Evolution of MLLMs via Continuous Softened Retracing reSampling
Yunyao Yu, Zhengxian Wu, Zhuohong Chen +6
In the unsupervised self-evolution of Multimodal Large Language Models, the quality of feedback signals during post-training is pivotal for stable and effective learning. However,…
When Models Judge Themselves: Unsupervised Self-Evolution for Multimodal Reasoning
Zhengxian Wu, Kai Shi, Chuanrui Zhang +10
Recent progress in multimodal large language models has led to strong performance on reasoning tasks, but these improvements largely rely on high-quality annotated data or teacher-…
PSGait: Gait Recognition using Parsing Skeleton
Hangrui Xu, Zhengxian Wu, Chuanrui Zhang +4
Gait recognition has emerged as a robust biometric modality due to its non-intrusive nature. Conventional gait recognition methods mainly rely on silhouettes or skeletons. While ef…
Language-Guided and Motion-Aware Gait Representation for Generalizable Recognition
Zhengxian Wu, Chuanrui Zhang, Shenao Jiang +6
Gait recognition is emerging as a promising technology and an innovative field within computer vision, with a wide range of applications in remote human identification. However, ex…