52 citations · 63 across the 12 of their papers we have counts for
7 papers · 1 filter
SRD-GUARD: A Defense Framework of LLMs via Semantic Rewriting and Joint Multi-Model Scoring for Latent Intent Exposure
Qi Wang, Chengcheng Wan, Jiangtao Wang
Large language models (LLMs) are increasingly deployed in safety-critical applications, yet jailbreak attacks can conceal harmful intent through role-playing, fictional scenarios,…
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction
Yuling Shi, Zhensu Sun, Junsen Dong +3
Evaluating LLM agents is essential for guiding their development, yet it has grown prohibitively expensive: a single pass of a frontier model over an agentic benchmark can cost hun…
Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?
Silin Chen, Yufei Yang, Xiaodong Gu +3
Repository-level coding benchmarks have become the standard for evaluating coding agents, yet they inherently suffer from data leakage because they are built upon popular open-sour…
Automated jailbreak attack targeting multiple defense strategies
Qi Wang, Chengcheng Wan, Weijia He +4
Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks. However, their safety remains a critical concern due to their susceptibility to…
MELD: Multi-Task Equilibrated Learning Detector for AI-Generated Text
Chenjun Li, Cheng Wan, Johannes C. Paetzold
Large language models are now embedded in everyday writing workflows, making reliable AI-generated text detection important for academic integrity, content moderation, and provenan…
EffiSkill: Agent Skill Based Automated Code Efficiency Optimization
Zimu Wang, Yuling Shi, Mengfan Li +4
Code efficiency is a fundamental aspect of software quality, yet how to harness large language models (LLMs) to optimize programs remains challenging. Prior approaches have sought…