8 papers
Exploring the Secondary Risks of Large Language Models
Jiawei Chen, Zhengwei Fang, Yu Tian +4
Ensuring the safety and alignment of Large Language Models is a significant challenge with their growing integration into critical applications and societal functions. While prior…
Reasoning as State Transition: A Representational Analysis of Reasoning Evolution in Large Language Models
Siyuan Zhang, Jialian Li, Yichi Zhang +3
Large Language Models have achieved remarkable performance on reasoning tasks, motivating research into how this ability evolves during training. Prior work has primarily analyzed…
KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
Shuyuan Liu, Jiawei Chen, Xiao Yang +2
With the widespread application of large language models (LLMs) in various fields, the security challenges they face have become increasingly prominent, especially the issue of jai…
From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models
Yuying Shang, Xinyi Zeng, Yutao Zhu +6
Hallucinations in large vision-language models (LVLMs) are a significant challenge, i.e., generating objects that are not presented in the visual input, which impairs their reliabi…
A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents
Hang Su, Jun Luo, Chang Liu +4
Recent advances in large language models (LLMs) have catalyzed the rise of autonomous AI agents capable of perceiving, reasoning, and acting in dynamic, open-ended environments. Th…
STAIR: Improving Safety Alignment with Introspective Reasoning
Yichi Zhang, Siyuan Zhang, Yao Huang +7
Ensuring the safety and harmlessness of Large Language Models (LLMs) has become equally critical as their performance in applications. However, existing safety alignment methods ty…