2 papers
cs.CR2026
LLM-Based Penetration Testing in the Presence of Honeypots
Xinhong Xie, Piyush Nagasubramaniam, Neeraj Karamchandani +1
Large language model (LLM) agents are increasingly employed for offensive cybersecurity tasks such as automated vulnerability discovery, reconnaissance, and penetration testing. Th…
cs.CL2026
Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit
Shuyi Fan, Boyuan Deng, Mengyu Xu +3
Audits of LLM judges certify a bias by contrasting matched conditions, and the strongest designs difference twice: a within-item contrast between two candidate responses, differenc…