activity
20242026
collaborators

6 papers

cs.AI2026

Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents

Muyu He, Anand Kumar, Tsach Mackey +3

Despite rapid progress in building conversational AI agents, robustness is still largely untested. Small shifts in user behavior, such as being more impatient, incoherent, or skept…

cs.CL2026

: Benchmarking AI Agents for Long-Term Planning and Consistent Execution

Muyu He, Adit Jain, Anand Kumar +4

As LLM agents tackle increasingly complex tasks, a critical question is whether they can maintain strategic coherence over long horizons: planning under uncertainty, learning from…

cs.CL2025

The Valley of Code Reasoning: Scaling Knowledge Distillation of Large Language Models

Muyu He, Muhammad Ali Shafique, Anand Kumar +2

Distilling the thinking traces of a Large Language Model (LLM) with reasoning capabilities into a smaller model has been proven effective. Yet, there is a scarcity of work done on…

cs.CL2025

Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models

Meghana Rajeev, Rajkumar Ramamurthy, Prapti Trivedi +5

We investigate the robustness of reasoning models trained for step-by-step problem solving by introducing query-agnostic adversarial triggers - short, irrelevant text that, when ap…

cs.CL2024

VERITAS: A Unified Approach to Reliability Evaluation

Rajkumar Ramamurthy, Meghana Arakkal Rajeev, Oliver Molenschot +2

Large language models (LLMs) often fail to synthesize information from their context to generate an accurate response. This renders them unreliable in knowledge intensive settings…

cs.CL2024

Self-rationalization improves LLM as a fine-grained judge

Prapti Trivedi, Aditya Gulati, Oliver Molenschot +7

LLM-as-a-judge models have been used for evaluating both human and AI generated content, specifically by providing scores and rationales. Rationales, in addition to increasing tran…