8 papers
SapiensID 2.0: Aligning Human Recognition Foundation Models with Human Perception
Yiyang Su, Jie Zhu, Feng Liu +2
While foundation models have significantly advanced human recognition across diverse modalities, they predominantly rely on static, geometric feature extraction. This approach fund…
Subtoken Vision Transformer for Fine-grained Recognition
Jie Zhu, Ivy Zhang, Minchul Kim +1
We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition. Standard Vision Transformers compress each fixed-size pa…
LocalScore: Local Density-Aware Similarity Scoring for Biometrics
Yiyang Su, Minchul Kim, Jie Zhu +4
Open-set biometrics faces challenges with probe subjects who may not be enrolled in the gallery, as traditional biometric systems struggle to detect these non-mated probes. Despite…
On the Holistic Approach for Detecting Human Image Forgery
Xiao Guo, Jie Zhu, Anil Jain +1
The rapid advancement of AI-generated content (AIGC) has escalated the threat of deepfakes, from facial manipulations to the synthesis of entire photorealistic human bodies. Howeve…
Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?
Jie Zhu, Yiyang Su, Xiaoming Liu
Multi-modal large language models (MLLMs) exhibit strong general-purpose capabilities, yet still struggle on Fine-Grained Visual Classification (FGVC), a core perception task that…
GEP: A GCG-Based method for extracting personally identifiable information from chatbots built on small language models
Jieli Zhu, Vi Ngoc-Nha Tran
Small language models (SLMs) become unprecedentedly appealing due to their approximately equivalent performance compared to large language models (LLMs) in certain fields with less…