3 citations · 7 across the 41 of their papers we have counts for
19 papers · 1 filter
CIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model Agents
Jiancheng Wang, Mingli Zhu, Tong Zhang +4
Visual world-model agents such as DreamerV3 act through a recurrent latent state rather than a single observation, which weakens frame-wise observation attacks and makes their pert…
SafeCA: Safe Cross-Attention Localization and Regulation for Text-to-Video Jailbreak Defense
Siyuan Liang, Yupeng Qiu, Junfeng Fang +3
Text-to-Video (T2V) generative models are vulnerable to jailbreak attacks in real-world deployment, leading them to produce harmful or inappropriate content. Existing defense appro…
Benchmarking the Robustness of Autonomous Driving to Environmental Illusions: A Lane Perception Perspective
Tianyuan Zhang, Xianglong Liu, Aishan Liu +6
Environmental illusions (eg., shadows, reflections, and tire marks) are naturally existing yet overlooked phenomena in real-world driving environments. They can disturb visual perc…
Who Generated This 3D Asset? Learning Source Attribution for Generative 3D Models
Sihan Ma, Siyuan Liang, Dacheng Tao
Generative 3D models are deployed in gaming, robotics, and immersive creation, making source attribution critical: given a 3D asset, can we identify whether and which generative mo…
CtrlAttack: A Unified Attack on World-Model Control in Diffusion Models
Shuhan Xu, Siyuan Liang, Hongling Zheng +4
Diffusion-based image-to-video (I2V) models increasingly exhibit world-model-like properties by implicitly capturing temporal dynamics. However, existing studies have mainly focuse…
Review of Hallucination Understanding in Large Language and Vision Models
Zhengyi Ho, Siyuan Liang, Dacheng Tao
The widespread adoption of large language and vision models in real-world applications has made urgent the need to address hallucinations -- instances where models produce incorrec…