2 papers
cs.SE2026
CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding
Hanjun Luo, Chiming Ni, Jiaheng Wen +9
LLM-powered coding agents are reshaping the development paradigm. However, existing evaluation systems, neither traditional tests for humans nor benchmarks for LLMs, fail to captur…
cs.AI2026
AgentAuditor: Human-Level Safety and Security Evaluation for LLM Agents
Hanjun Luo, Shenyu Dai, Chiming Ni +5
Despite the rapid advancement of LLM-based agents, the reliable evaluation of their safety and security remains a significant challenge. Existing rule-based or LLM-based evaluators…