activity
20242026
collaborators

5 papers

cs.CV2026

Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception

Yanpeng Sun, Jing Hao, Ke Zhu +6

Training Large Multimodality Models (LMMs) relies on descriptive image caption that connects image and language. Existing methods for generating such captions often rely on distill…

cs.CV2025

Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models

Guosheng Zhang, Keyao Wang, Haixiao Yue +5

Face Anti-Spoofing (FAS) is essential for ensuring the security and reliability of facial recognition systems. Most existing FAS methods are formulated as binary classification tas…

cs.CV2024

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities

Huan Liu, Lingyu Xiao, Jiangjiang Liu +4

With the rapid advancement of Multimodal Large Language Models (MLLMs), a variety of benchmarks have been introduced to evaluate their capabilities. While most evaluations have foc…

cs.CV2024

ALoRE: Efficient Visual Adaptation via Aggregating Low Rank Experts

Sinan Du, Guosheng Zhang, Keyao Wang +7

Parameter-efficient transfer learning (PETL) has become a promising paradigm for adapting large-scale vision foundation models to downstream tasks. Typical methods primarily levera…

cs.LG2024

Continual SFT Matches Multimodal RLHF with Negative Supervision

Ke Zhu, Yu Wang, Yanpeng Sun +4

Multimodal RLHF usually happens after supervised finetuning (SFT) stage to continually improve vision-language models' (VLMs) comprehension. Conventional wisdom holds its superiori…