6 papers
Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents
Muyu He, Anand Kumar, Tsach Mackey +3
Despite rapid progress in building conversational AI agents, robustness is still largely untested. Small shifts in user behavior, such as being more impatient, incoherent, or skept…
: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
Muyu He, Adit Jain, Anand Kumar +4
As LLM agents tackle increasingly complex tasks, a critical question is whether they can maintain strategic coherence over long horizons: planning under uncertainty, learning from…
The Valley of Code Reasoning: Scaling Knowledge Distillation of Large Language Models
Muyu He, Muhammad Ali Shafique, Anand Kumar +2
Distilling the thinking traces of a Large Language Model (LLM) with reasoning capabilities into a smaller model has been proven effective. Yet, there is a scarcity of work done on…
Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models
Meghana Rajeev, Rajkumar Ramamurthy, Prapti Trivedi +5
We investigate the robustness of reasoning models trained for step-by-step problem solving by introducing query-agnostic adversarial triggers - short, irrelevant text that, when ap…
VERITAS: A Unified Approach to Reliability Evaluation
Rajkumar Ramamurthy, Meghana Arakkal Rajeev, Oliver Molenschot +2
Large language models (LLMs) often fail to synthesize information from their context to generate an accurate response. This renders them unreliable in knowledge intensive settings…
Self-rationalization improves LLM as a fine-grained judge
Prapti Trivedi, Aditya Gulati, Oliver Molenschot +7
LLM-as-a-judge models have been used for evaluating both human and AI generated content, specifically by providing scores and rationales. Rationales, in addition to increasing tran…