activity
20242026
collaborators

8 papers

cs.LG2026

Exploring the Secondary Risks of Large Language Models

Jiawei Chen, Zhengwei Fang, Yu Tian +4

Ensuring the safety and alignment of Large Language Models is a significant challenge with their growing integration into critical applications and societal functions. While prior…

cs.CL2026

Reasoning as State Transition: A Representational Analysis of Reasoning Evolution in Large Language Models

Siyuan Zhang, Jialian Li, Yichi Zhang +3

Large Language Models have achieved remarkable performance on reasoning tasks, motivating research into how this ability evolves during training. Prior work has primarily analyzed…

cs.CR2025

KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs

Shuyuan Liu, Jiawei Chen, Xiao Yang +2

With the widespread application of large language models (LLMs) in various fields, the security challenges they face have become increasingly prominent, especially the issue of jai…

cs.CL2025

From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models

Yuying Shang, Xinyi Zeng, Yutao Zhu +6

Hallucinations in large vision-language models (LVLMs) are a significant challenge, i.e., generating objects that are not presented in the visual input, which impairs their reliabi…

cs.AI2025

A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents

Hang Su, Jun Luo, Chang Liu +4

Recent advances in large language models (LLMs) have catalyzed the rise of autonomous AI agents capable of perceiving, reasoning, and acting in dynamic, open-ended environments. Th…

cs.CL2025

STAIR: Improving Safety Alignment with Introspective Reasoning

Yichi Zhang, Siyuan Zhang, Yao Huang +7

Ensuring the safety and harmlessness of Large Language Models (LLMs) has become equally critical as their performance in applications. However, existing safety alignment methods ty…