3 papers
cs.AI2026
A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation
Rahul Khedar, Eshita, Sneha Teja Sree Reddy Thondapu +6
Evaluating tool-augmented LLM agents requires diverse, realistic user inputs yet most evaluation frameworks use flat role descriptions ("you are an angry customer") that produce ne…
cs.AI2026
AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval
Gunja Agarwal, Arup Kumar Das, Arun Menon +2
Evaluation of agentic information retrieval remains limited to scripted interactions with uniform users, missing both natural personality diversity and adversarial brittleness. We…
cs.AI2026
State-Grounded Multi-Agent Synthetic Data Generation for Tool-Augmented LLMs
Rahul Khedar, Eshita, Sneha Teja Sree Reddy Thondapu +10
Training tool-augmented LLM agents requires large corpora of multi-turn, tool-grounded conversational data that is expensive to annotate, privacy-constrained in production settings…