4 papers
EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses
Tanmay Sah, Dolly Sah, Harshul Jain +1
LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution harnesses at runtime. Such self-evolution can improve capability, but a successful mut…
The AI Evaluability Gap: The Missing Layer for Managing Risk and Sustaining Value
Vishal Srivastava, Tanmay Sah
Organizations deploying AI face two fundamental governance challenges: managing AI risk and sustaining AI value. Both depend on evidence whose sufficiency cannot be taken for grant…
The Verifier Tax: Horizon Dependent Safety Success Tradeoffs in Tool Using LLM Agents
Tanmay Sah, Vishal Srivastava, Dolly Sah +1
We study how runtime enforcement against unsafe actions affects end-to-end task performance in multi-step tool using large language model (LLM) agents. Using tau-bench across Airli…
Quantifying Automation Risk in High-Automation AI Systems: A Bayesian Framework for Failure Propagation and Optimal Oversight
Vishal Srivastava, Tanmay Sah
Organizations across finance, healthcare, transportation, content moderation, and critical infrastructure are rapidly deploying highly automated AI systems, yet they lack principle…