3 papers
cs.AI2024
Autonomous Evaluation of LLMs for Truth Maintenance and Reasoning Tasks
Rushang Karia, Daniel Bramblett, Daksh Dobhal +1
This paper presents AutoEval, a novel benchmark for scaling Large Language Model (LLM) assessment in formal tasks with clear notions of correctness, such as truth maintenance in tr…
cs.RO2024
Using Explainable AI and Hierarchical Planning for Outreach with Robots
Rushang Karia, Jayesh Nagpal, Daksh Dobhal +4
Understanding how robots plan and execute tasks is crucial in today's world, where they are becoming more prevalent in our daily lives. However, teaching non-experts, such as K-12…
cs.CL2024
utoval: Autonomous Assessment of LLMs in Formal Synthesis and Interpretation Tasks
Rushang Karia, Daniel Bramblett, Daksh Dobhal +2
This paper presents utoval, a new approach for scaling LLM assessment in translating formal syntax -- such as first-order logic, regular expressions, etc -- to na…