papers

Publications (185)

cs.LG2022

Towards Adversarial Attack on Vision-Language Pre-training Models

Jiaming Zhang, Qi Yi, Jitao Sang

While vision-language pre-training model (VLP) has shown revolutionary improvements on various vision-language (V+L) tasks, the studies regarding its adversarial robustness remain…

cs.CV2026

NAP-Tuning: Neural Augmented Prompt Tuning for Adversarially Robust Vision-Language Models

Jiaming Zhang, Xin Wang, Xingjun Ma +3

Vision-Language Models (VLMs) such as CLIP have demonstrated remarkable capabilities in understanding relationships between visual and textual data through joint embedding spaces.…

cond-mat.mtrl-sci2012

Growth process and crystallographic properties of ammonia-induced vaterite

Qiaona Hu, Jiaming Zhang, Henry Teng +1

Metastable vaterite crystals were synthesized by increasing the pH and consequently the saturation states of Ca and CO3 containing solutions using an ammonia diffusion method. SEM…

cs.CV2024

OneBEV: Using One Panoramic Image for Bird's-Eye-View Semantic Mapping

Jiale Wei, Junwei Zheng, Ruiping Liu +3

In the field of autonomous driving, Bird's-Eye-View (BEV) perception has attracted increasing attention in the community since it provides more comprehensive information compared w…

cs.CV2026

Disrupting Hierarchical Reasoning: Adversarial Protection for Geographic Privacy in Multimodal Reasoning Models

Jiaming Zhang, Che Wang, Yang Cao +2

Multi-modal large reasoning models (MLRMs) pose significant privacy risks by inferring precise geographic locations from personal images through hierarchical chain-of-thought reaso…

cs.AI2026

GUITester: Enabling GUI Agents for Exploratory Defect Discovery

Yifei Gao, Jiang Wu, Xiaoyi Chen +5

Exploratory GUI testing is essential for software quality but suffers from high manual costs. While Multi-modal Large Language Model (MLLM) agents excel in navigation, they fail to…