6 papers
The Benchmark Illusion: Pruned LLMs Can Pass Multiple Choice but Fail to Answer
Rui Wen, Lu Sun, Jiayang Liu +3
Compressing large language models reduces memory use and inference cost, but it can also create failures that standard benchmarks miss. A pruned model may still perform well on mul…
When Memory Becomes a Vulnerability: Towards Multi-turn Jailbreak Attacks against Text-to-Image Generation Systems
Shiqian Zhao, Jiayang Liu, Yiming Li +9
Modern text-to-image (T2I) generation systems (e.g., DALLE 3) exploit the memory mechanism, which captures key information in multi-turn interactions for faithful generation…
RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
Runwei Guan, Rongsheng Hu, Shangshu Chen +10
Current roadside perception systems mainly focus on instance-level perception, which fall short in enabling interaction via natural language and reasoning about traffic behaviors i…
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
Youze Wang, Zijun Chen, Ruoyu Chen +8
Recent advancements in multimodal large language models for video understanding (videoLLMs) have enhanced their capacity to process complex spatiotemporal data. However, challenges…
T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak Attacks
Jiayang Liu, Siyuan Liang, Shiqian Zhao +5
In recent years, fueled by the rapid advancement of diffusion models, text-to-video (T2V) generation models have achieved remarkable progress, with notable examples including Pika,…
T2VShield: Model-Agnostic Jailbreak Defense for Text-to-Video Models
Siyuan Liang, Jiayang Liu, Jiecheng Zhai +5
The rapid development of generative artificial intelligence has made text to video models essential for building future multimodal world simulators. However, these models remain vu…