2 papers
cs.AI2026
No Universal Signal Predicts Sample-Level LLM Regression under Version Updates
Jia Sheng, Yiwei Lu
Frontier LLMs are updated frequently and typically outperform their predecessors in aggregate. But aggregate gains say little about individual samples: an update can still cause sa…
cs.CR2025
Evaluating LLM Generated Detection Rules in Cybersecurity
Anna Bertiger, Bobby Filar, Aryan Luthra +4
LLMs are increasingly pervasive in the security environment, with limited measures of their effectiveness, which limits trust and usefulness to security practitioners. Here, we pre…