3 papers
cs.AI2026
A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations
Hasan Najib Mahmud, Shreya Gupta, Isha Chaudhary +4
AI code agents are increasingly deployed to resolve real software issues, yet their reliability under superficial code variations remains poorly understood. We evaluate whether cod…
cs.PL2025
Lumos: Let there be Language Model System Certification
Isha Chaudhary, Vedaant Jain, Prineet Parhar +4
We introduce the first principled framework, Lumos, for specifying and formally certifying Language Model System (LMS) behaviors. Lumos is an imperative probabilistic programming D…
cs.AI2025
How Catastrophic is Your LLM? Certifying Risk in Conversation
Chengxiao Wang, Isha Chaudhary, Qian Hu +3
Large Language Models (LLMs) can produce catastrophic responses in conversational settings that pose serious risks to public safety and security. Existing evaluations often fail to…