Publications (185)
Towards Adversarial Attack on Vision-Language Pre-training Models
Jiaming Zhang, Qi Yi, Jitao Sang
While vision-language pre-training model (VLP) has shown revolutionary improvements on various vision-language (V+L) tasks, the studies regarding its adversarial robustness remain…
NAP-Tuning: Neural Augmented Prompt Tuning for Adversarially Robust Vision-Language Models
Jiaming Zhang, Xin Wang, Xingjun Ma +3
Vision-Language Models (VLMs) such as CLIP have demonstrated remarkable capabilities in understanding relationships between visual and textual data through joint embedding spaces.…
Growth process and crystallographic properties of ammonia-induced vaterite
Qiaona Hu, Jiaming Zhang, Henry Teng +1
Metastable vaterite crystals were synthesized by increasing the pH and consequently the saturation states of Ca and CO3 containing solutions using an ammonia diffusion method. SEM…
OneBEV: Using One Panoramic Image for Bird's-Eye-View Semantic Mapping
Jiale Wei, Junwei Zheng, Ruiping Liu +3
In the field of autonomous driving, Bird's-Eye-View (BEV) perception has attracted increasing attention in the community since it provides more comprehensive information compared w…
Disrupting Hierarchical Reasoning: Adversarial Protection for Geographic Privacy in Multimodal Reasoning Models
Jiaming Zhang, Che Wang, Yang Cao +2
Multi-modal large reasoning models (MLRMs) pose significant privacy risks by inferring precise geographic locations from personal images through hierarchical chain-of-thought reaso…
GUITester: Enabling GUI Agents for Exploratory Defect Discovery
Yifei Gao, Jiang Wu, Xiaoyi Chen +5
Exploratory GUI testing is essential for software quality but suffers from high manual costs. While Multi-modal Large Language Model (MLLM) agents excel in navigation, they fail to…