papers

Publications (7)

cs.CL2025

Direct Behavior Optimization: Unlocking the Potential of Lightweight LLMs

Hongming Yang, Shi Lin, Jun Shao +4

Lightweight Large Language Models (LwLLMs) are reduced-parameter, optimized models designed to run efficiently on consumer-grade hardware, offering significant advantages in resour…

cs.CR2025

LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models

Shi Lin, Hongming Yang, Rongchang Li +4

The rapid development of Large Language Models (LLMs) has brought impressive advancements across various tasks. However, despite these achievements, LLMs still pose inherent safety…

cs.CR2026

Before Agents Speak: Pre-hoc Failure Risk Inference in Multi-Agent Systems

Shi Lin, Chenpei Wang, Peng Qian +4

The paper introduces HalluProp, a framework that predicts which agents in a large‑language‑model based multi‑agent system are likely to hallucinate and estimates the overall system…

#multi-agent systems#hallucination detection#pre-hoc risk inference#LLM safety
cs.CR2025

A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasures

Dezhang Kong, Shi Lin, Zhenhua Xu +16

In recent years, Large-Language-Model-driven AI agents have exhibited unprecedented intelligence and adaptability. Nowadays, agents are undergoing a new round of evolution. They no…

cs.LG2026

Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions

Shi Lin, Peng Qian, Dinghao Liu +5

The paper introduces Recast, a framework that predicts safety risks in multi‑turn interactions with large language models by forecasting how risks evolve over dialogue trajectories…

#llm safety#risk forecasting#multi-turn dialogue#trajectory-level prediction
cs.CR2026

Web Fraud Attacks Against LLM-Driven Multi-Agent Systems

Dezhang Kong, Hujin Peng, Yilun Zhang +5

With the proliferation of LLM-driven multi-agent systems (MAS), the security of Web links has become a critical concern. Once MAS is induced to trust a malicious link, attackers ca…

cs.CR2025

Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025

Zonghao Ying, Siyang Wu, Run Hao +44

Multimodal Large Language Models (MLLMs) have enabled transformative advancements across diverse applications but remain susceptible to safety threats, especially jailbreak attacks…