2 papers
cs.CR2026
SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response
Lehan Wang, Boli Chen, Ruixue Ding +7
The paper presents SecRespond, a benchmark that evaluates large language model agents on post-compromise incident‑response tasks using forensic disk snapshots, alerts, and vulnerab…
cs.LG2026
ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking
Qiang Zhang, Boli Chen, Fanrui Zhang +14
Reinforcement learning has substantially improved the performance of LLM agents on tasks with verifiable outcomes, but it still struggles on open-ended agent tasks with vast soluti…