2 papers
cs.CR2026
LLM agents security duality: a comprehensive survey of self-security and empowered cybersecurity
Yiwei Xu, Yong Zhuang, Xuanming Liu +6
Large language model (LLM) agents are rapidly being integrated into real-world systems. Their autonomy and tool-use capabilities generate substantial value while simultaneously exp…
cs.AI2026
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…