3 papers
cs.AI2026
Monte Carlo Query Search: Active Capability Assessment of AI Agents
Daniel Bramblett, Rushang Karia, Adrian Ciotinga +3
Black-box AI (BBAI) systems, including foundation-model agents, are increasingly used for sequential decision making. Safe deployment requires methods for characterizing what such…
cs.RO2025
Using Explainable AI and Hierarchical Planning for Outreach with Robots
Rushang Karia, Jayesh Nagpal, Daksh Dobhal +4
Understanding how robots plan and execute tasks is crucial in today's world, where they are becoming more prevalent in our daily lives. However, teaching non-experts, such as K-12…
cs.AI2025
Autonomous Evaluation of LLMs for Truth Maintenance and Reasoning Tasks
Rushang Karia, Daniel Bramblett, Daksh Dobhal +1
This paper presents AutoEval, a novel benchmark for scaling Large Language Model (LLM) assessment in formal tasks with clear notions of correctness, such as truth maintenance in tr…