9 papers
RISA: Response Inspection and Selective Actions for Refusal Calibration in Large Language Models
Wenhan Chang, Tianqing Zhu, Ping Xiong +2
Reliable refusal behavior requires Large Language Models (LLMs) to reject harmful prompts with only answering benign ones. Incorrect refusal behavior can either expose users to har…
From "What-If" to "What-Is": Counterfactual Thinking-Inspired Semantic Alignment for Visual Brain Decoding
Kaitao Yan, Chi Liu, Congcong Zhu +5
Visual brain decoding reconstructs visual content perceived by a person from neural measurements such as fMRI, providing a computational approach to studying how visual information…
Auditing Machine Unlearning: A Systematic Research on Whether Models Truly Forget
Dayong Ye, Tianqing Zhu, Ruiding Huang +5
Machine unlearning has been extensively studied in response to growing privacy concerns and regulatory requirements. However, auditing whether unlearning algorithms have truly eras…
Seeing No Evil: Blinding Large Vision-Language Models to Safety Instructions via Adversarial Attention Hijacking
Jingru Li, Wei Ren, Tianqing Zhu
Large Vision-Language Models (LVLMs) rely on attention-based retrieval of safety instructions to maintain alignment during generation. Existing attacks typically optimize image per…
Unreal Thinking: Chain-of-Thought Hijacking via Two-stage Backdoor
Wenhan Chang, Tianqing Zhu, Ping Xiong +2
Large Language Models (LLMs) are increasingly deployed in settings where Chain-of-Thought (CoT) is interpreted by users. This creates a new safety risk: attackers may manipulate th…
Are LLMs Ready for Computer Science Education? A Cross-Domain, Cross-Lingual and Cognitive-Level Evaluation Using Professional Certification Exams
Chen Gao, Chi Liu, Zhengquan Luo +10
Large language models (LLMs) are increasingly applied in computer science education for tasks such as tutoring, content generation, and code assessment. However, systematic evaluat…