collaborators

6 papers

cs.CL2026

The Benchmark Illusion: Pruned LLMs Can Pass Multiple Choice but Fail to Answer

Rui Wen, Lu Sun, Jiayang Liu +3

Compressing large language models reduces memory use and inference cost, but it can also create failures that standard benchmarks miss. A pruned model may still perform well on mul…

cs.CV2026

When Memory Becomes a Vulnerability: Towards Multi-turn Jailbreak Attacks against Text-to-Image Generation Systems

Shiqian Zhao, Jiayang Liu, Yiming Li +9

Modern text-to-image (T2I) generation systems (e.g., DALLE 3) exploit the memory mechanism, which captures key information in multi-turn interactions for faithful generation…

cs.CV2025

RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System

Runwei Guan, Rongsheng Hu, Shangshu Chen +10

Current roadside perception systems mainly focus on instance-level perception, which fall short in enabling interaction via natural language and reasoning about traffic behaviors i…

cs.CV2025

Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding

Youze Wang, Zijun Chen, Ruoyu Chen +8

Recent advancements in multimodal large language models for video understanding (videoLLMs) have enhanced their capacity to process complex spatiotemporal data. However, challenges…

cs.CV2025

T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak Attacks

Jiayang Liu, Siyuan Liang, Shiqian Zhao +5

In recent years, fueled by the rapid advancement of diffusion models, text-to-video (T2V) generation models have achieved remarkable progress, with notable examples including Pika,…

cs.CR2025

T2VShield: Model-Agnostic Jailbreak Defense for Text-to-Video Models

Siyuan Liang, Jiayang Liu, Jiecheng Zhai +5

The rapid development of generative artificial intelligence has made text to video models essential for building future multimodal world simulators. However, these models remain vu…