5 papers
Indi-RomCoM: Code-Mixed Benchmark for Evaluating LLMs on Romanized Indic-English Instructions
Avisha Das, Mihir Parmar, Mohana Ramnath +1
Romanized Code Mixing (RCM), where bilingual speakers fluidly blend local languages with English in Roman script, has emerged as the dominant form of communication across multiling…
Monte Carlo Query Search: Active Capability Assessment of AI Agents
Daniel Bramblett, Rushang Karia, Adrian Ciotinga +3
Black-box AI (BBAI) systems, including foundation-model agents, are increasingly used for sequential decision making. Safe deployment requires methods for characterizing what such…
From Real World to Logic and Back: Learning Generalizable Relational Concepts For Long Horizon Robot Planning
Naman Shah, Jayesh Nagpal, Siddharth Srivastava
Robots still lag behind humans in their ability to generalize from limited experience, particularly when transferring learned behaviors to long-horizon tasks in unseen environments…
Using Explainable AI and Hierarchical Planning for Outreach with Robots
Rushang Karia, Jayesh Nagpal, Daksh Dobhal +4
Understanding how robots plan and execute tasks is crucial in today's world, where they are becoming more prevalent in our daily lives. However, teaching non-experts, such as K-12…
AI Planning: A Primer and Survey (Preliminary Report)
Dillon Z. Chen, Pulkit Verma, Siddharth Srivastava +2
Automated decision-making is a fundamental topic that spans multiple sub-disciplines in AI: reinforcement learning (RL), AI planning (AP), foundation models, and operations researc…