collaborators

10 papers

cs.CV2026

SapiensID 2.0: Aligning Human Recognition Foundation Models with Human Perception

Yiyang Su, Jie Zhu, Feng Liu +2

While foundation models have significantly advanced human recognition across diverse modalities, they predominantly rely on static, geometric feature extraction. This approach fund…

cs.CV2026

Subtoken Vision Transformer for Fine-grained Recognition

Jie Zhu, Ivy Zhang, Minchul Kim +1

We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition. Standard Vision Transformers compress each fixed-size pa…

cs.CV2026

Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?

Jie Zhu, Yiyang Su, Xiaoming Liu

Multi-modal large language models (MLLMs) exhibit strong general-purpose capabilities, yet still struggle on Fine-Grained Visual Classification (FGVC), a core perception task that…

cs.CV2026

FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human Recognition

Jie Zhu, Xiao Guo, Yiyang Su +2

Model fusion is a key strategy for robust recognition in unconstrained scenarios, as different models provide complementary strengths. This is especially important for whole-body h…

cs.CV2026

Interpretable Perception and Reasoning for Audiovisual Geolocation

Yiyang Su, Xiaoming Liu

While recent advances in Multimodal Large Language Models (MLLMs) have improved image-based localization, precise global geolocation remains a formidable challenge due to the inher…

cs.CV2026

LocalScore: Local Density-Aware Similarity Scoring for Biometrics

Yiyang Su, Minchul Kim, Jie Zhu +4

Open-set biometrics faces challenges with probe subjects who may not be enrolled in the gallery, as traditional biometric systems struggle to detect these non-mated probes. Despite…