11 papers
Adversarial Attacks on Deep OCR Systems
Wenbo Sun, Hongzong LI, Yanyun Wang +5
Deep-OCR (DeepSeek-OCR) advances document recognition by treating the visual modality as an optical compression medium, enabling long-context OCR at low token cost. However, its in…
Acoustic Interference: A New Paradigm Weaponizing Acoustic Latent Semantic for Universal Jailbreak against Large Audio Language Models
Yanyun Wang, Yu Huang, Zi Liang +2
The integration of audio modality into Large Audio Language Models (LALMs) significantly expands their attack surface. Existing jailbreak paradigms predominantly treat audio as a c…
PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation
Tianxin Xie, Wentao Lei, Kai Jiang +27
Text-to-audio-video (T2AV) generation is central to applications such as filmmaking and world modeling. However, current models often fail to produce physically plausible sounds. P…
Can a Single Message Paralyze the AI Infrastructure? The Rise of AbO-DDoS Attacks through Targeted Mobius Injection
Zi Liang, Ronghua Li, Yanyun Wang +2
Large Language Model (LLM) agents have emerged as key intermediaries, orchestrating complex interactions between human users and a wide range of digital services and LLM infrastruc…
Robust Alignment: Harmonizing Clean Accuracy and Adversarial Robustness in Adversarial Training
Yanyun Wang, Qingqing Ye, Li Liu +2
Adversarial Training (AT) is one of the most effective methods for developing robust deep neural networks (DNNs). However, AT faces a trade-off problem between clean accuracy and a…
Revitalizing Canonical Pre-Alignment for Irregular Multivariate Time Series Forecasting
Ziyu Zhou, Yiming Huang, Yanyun Wang +3
Irregular multivariate time series (IMTS), characterized by uneven sampling and inter-variate asynchrony, fuel many forecasting applications yet remain challenging to model efficie…