2 papers
cs.CR2026
RvB: Automating AI System Hardening via Iterative Red-Blue Games
Lige Huang, Zicheng Liu, Jie Zhang +3
The dual offensive and defensive utility of Large Language Models (LLMs) highlights a critical gap in AI security: the lack of unified frameworks for dynamic, iterative adversarial…
cs.CR2025
PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
Zicheng Liu, Lige Huang, Jie Zhang +3
The increasing autonomy of Large Language Models (LLMs) necessitates a rigorous evaluation of their potential to aid in cyber offense. Existing benchmarks often lack real-world com…