activity
20242026
most citedHybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference

1 citations · 2 across the 20 of their papers we have counts for

collaborators

22 papers

cs.AI2026

JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills

Xiaoyu Wen, Jiajia Li, Zhida He +11

Automated red-teaming has produced a growing collection of attack strategies, yet they typically remain scattered across prompts and workflows, making them difficult to systematica…

cs.LG2026

ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions

Xinzhe Huang, Biwu Yao, Kedong Xiu +4

Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formulate safety assessment as a de…

cs.CR2026

AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents

Ruoyu Wang, Heng Zhao, Renjie Wu +4

Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by tools. This dependence allows…

cs.AI2026

REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment

Zhengze Huang, Luyang Yu, Di Hong +5

Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment. Hallucinations in LRMs arise from two distinc…

cs.CL2026

An Early Warning of Emerging Biosecurity Risks in Frontier LLMs

Zhida He, Xia Hu, Baichen Le +20

Frontier large language models (LLMs) are increasingly integrated into scientific workflows, yet their growing biological capabilities may outpace current safeguards. To assess the…

cs.CR2026

Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning

Yuxuan Huang, Xingyu Zeng, Tianhang Zheng +1

Released aligned large language models remain vulnerable to malicious downstream finetuning. Existing defenses are largely designed for the fine-tuning-as-a-service (FTaaS) paradig…