activity
20242026
collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors

Shuhaib Mehri, Philippe Laban, Sumuk Shashidhar +4

As user simulators are increasingly used for interactive training and evaluation of AI assistants, it is essential that they represent the diverse behaviors of real users. While ex…

cs.CL2025

AURA: A Diagnostic Framework for Tracking User Satisfaction of Interactive Planning Agents

Takyoung Kim, Janvijay Singh, Shuhaib Mehri +6

The growing capabilities of large language models (LLMs) in instruction-following and context-understanding lead to the era of agents with numerous applications. Among these, task…

cs.CL2025

Question Generation for Assessing Early Literacy Reading Comprehension

Xiaocheng Yang, Sumuk Shashidhar, Dilek Hakkani-Tur

Assessment of reading comprehension through content-based interactions plays an important role in the reading acquisition process. In this paper, we propose a novel approach for ge…

cs.CL2025

YourBench: Easy Custom Evaluation Sets for Everyone

Sumuk Shashidhar, Clémentine Fourrier, Alina Lozovskia +3

Evaluating large language models (LLMs) effectively remains a critical bottleneck, as traditional static benchmarks suffer from saturation and contamination, while human evaluation…

cs.CL2024

Unsupervised Human Preference Learning

Sumuk Shashidhar, Abhinav Chinta, Vaibhav Sahai +1

Large language models demonstrate impressive reasoning abilities but struggle to provide personalized content due to their lack of individual user preference information. Existing…