3 papers
cs.LG2026
ALERT: Zero-shot LLM Jailbreak Detection via Internal Discrepancy Amplification
Xiao Lin, Philip Li, Zhichen Zeng +6
Despite rich safety alignment strategies, large language models (LLMs) remain highly susceptible to jailbreak attacks, which compromise safety guardrails and pose serious security…
cs.CL2025
Taming Knowledge Conflicts in Language Models
Gaotang Li, Yuzhong Chen, Hanghang Tong
Language Models (LMs) often encounter knowledge conflicts when parametric memory contradicts contextual knowledge. Previous works attribute this conflict to the interplay between "…
cs.LG2025
CATS: Mitigating Correlation Shift for Multivariate Time Series Classification
Xiao Lin, Zhichen Zeng, Tianxin Wei +3
Unsupervised Domain Adaptation (UDA) leverages labeled source data to train models for unlabeled target data. Given the prevalence of multivariate time series (MTS) data across var…