Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
NAAMSE: Framework for Evolutionary Security Evaluation of Agents
Kunal Pai, Parth Shah, Harshil Patel
AI agents are increasingly deployed in production, yet their security evaluations remain bottlenecked by manual red-teaming or static benchmarks that fail to model adaptive, multi-…
cs.AI2025
How Many Instructions Can LLMs Follow at Once?
Daniel Jaroslawicz, Brendan Whiting, Parth Shah +1
Production-grade LLM systems require robust adherence to dozens or even hundreds of instructions simultaneously. However, the instruction-following capabilities of LLMs at high ins…