1 citations · 1 across the 10 of their papers we have counts for
11 papers
HyperCLIP++: Fine-tuning CLIP forOpen-vocabulary Semantic Segmentation in Hyperbolic Space
Zelin Peng, Zhengqin Xu, Changsong Wen +4
CLIP, a foundational vision-language model, has emerged as a powerful tool for open-vocabulary semantic segmentation. While freezing CLIP's text encoder is known to preserve its ge…
Patch-Discontinuity Mining for Generalized Deepfake Detection
Huanhuan Yuan, Yang Ping, Zhengqin Xu +3
The rapid advancement of generative artificial intelligence has enabled the creation of highly realistic fake facial images, posing serious threats to personal privacy and the inte…
HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
Zelin Peng, Zhengqin Xu, Qingyang Liu +2
Multi-modal large language models (MLLMs) have emerged as a transformative approach for aligning visual and textual understanding. They typically require extremely high computation…
NEARL: Interacted Query Adaptation with Orthogonal Regularization for Medical Vision-Language Understanding
Zelin Peng, Yichen Zhao, Yu Huang +5
Computer-aided medical image analysis is crucial for disease diagnosis and treatment planning. While vision-language models (VLMs) such as CLIP exhibit strong generalization abilit…
Robust SAM: On the Adversarial Robustness of Vision Foundation Models
Jiahuan Long, Zhengqin Xu, Tingsong Jiang +4
The Segment Anything Model (SAM) is a widely used vision foundation model with diverse applications, including image segmentation, detection, and tracking. Given SAM's wide applica…
Parameter-Free Fine-tuning via Redundancy Elimination for Vision Foundation Models
Jiahuan Long, Tingsong Jiang, Wen Yao +5
Vision foundation models (VFMs) have demonstrated remarkable capabilities in learning universal visual representations. However, adapting these models to downstream tasks conventio…