7 papers
EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs
Liang Lin, Chunxi Luo, Kaiwen Luo +9
Audio Large Language Models (ALLMs) are highly vulnerable to real-world noise, which often induces severe semantic drift and hallucinations. Existing robustness methods primarily r…
Large Vision-Language Models Get Lost in Attention
Gongli Xi, Ye Tian, Mengyu Yang +5
Despite the rapid evolution of training paradigms, the decoder backbone of large vision--language models (LVLMs) remains fundamentally rooted in the residual-connection Transformer…
Packet-Level DDoS Data Augmentation Using Dual-Stream Temporal-Field Diffusion
Gongli Xi, Ye Tian, Yannan Hu +3
In response to Distributed Denial of Service (DDoS) attacks, recent research efforts increasingly rely on Machine Learning (ML)-based solutions, whose effectiveness largely depends…
SaFeR-ToolKit: Structured Reasoning via Virtual Tool Calling for Multimodal Safety
Zixuan Xu, Tiancheng He, Huahui Yi +7
Vision-language models remain susceptible to multimodal jailbreaks and over-refusal because safety hinges on both visual evidence and user intent, while many alignment pipelines su…
ClueTracer: Question-to-Vision Clue Tracing for Training-Free Hallucination Suppression in Multimodal Reasoning
Gongli Xi, Kun Wang, Zeming Gao +4
Large multimodal reasoning models solve challenging visual problems via explicit long-chain inference: they gather visual clues from images and decode clues into textual tokens. Ye…
CoPHo: Classifier-guided Conditional Topology Generation with Persistent Homology
Gongli Xi, Ye Tian, Mengyu Yang +5
The structure of topology underpins much of the research on performance and robustness, yet available topology data are typically scarce, necessitating the generation of synthetic…