9 papers
SapiensID 2.0: Aligning Human Recognition Foundation Models with Human Perception
Yiyang Su, Jie Zhu, Feng Liu +2
While foundation models have significantly advanced human recognition across diverse modalities, they predominantly rely on static, geometric feature extraction. This approach fund…
Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?
Jie Zhu, Yiyang Su, Xiaoming Liu
Multi-modal large language models (MLLMs) exhibit strong general-purpose capabilities, yet still struggle on Fine-Grained Visual Classification (FGVC), a core perception task that…
FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human Recognition
Jie Zhu, Xiao Guo, Yiyang Su +2
Model fusion is a key strategy for robust recognition in unconstrained scenarios, as different models provide complementary strengths. This is especially important for whole-body h…
Interpretable Perception and Reasoning for Audiovisual Geolocation
Yiyang Su, Xiaoming Liu
While recent advances in Multimodal Large Language Models (MLLMs) have improved image-based localization, precise global geolocation remains a formidable challenge due to the inher…
LocalScore: Local Density-Aware Similarity Scoring for Biometrics
Yiyang Su, Minchul Kim, Jie Zhu +4
Open-set biometrics faces challenges with probe subjects who may not be enrolled in the gallery, as traditional biometric systems struggle to detect these non-mated probes. Despite…
HAMoBE: Hierarchical and Adaptive Mixture of Biometric Experts for Video-based Person ReID
Yiyang Su, Yunping Shi, Feng Liu +1
Recently, research interest in person re-identification (ReID) has increasingly focused on video-based scenarios, which are essential for robust surveillance and security in varied…