7 papers
Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors
Shuhaib Mehri, Philippe Laban, Sumuk Shashidhar +4
As user simulators are increasingly used for interactive training and evaluation of AI assistants, it is essential that they represent the diverse behaviors of real users. While ex…
Combinatorial Creativity: A New Frontier in Generalization Abilities
Samuel Schapiro, Sumuk Shashidhar, Alexi Gladstone +4
Artificial intelligence (AI) systems, and Large Language Models (LLMs) in particular, are increasingly employed for creative tasks like scientific idea generation, constituting a f…
AURA: A Diagnostic Framework for Tracking User Satisfaction of Interactive Planning Agents
Takyoung Kim, Janvijay Singh, Shuhaib Mehri +6
The growing capabilities of large language models (LLMs) in instruction-following and context-understanding lead to the era of agents with numerous applications. Among these, task…
Question Generation for Assessing Early Literacy Reading Comprehension
Xiaocheng Yang, Sumuk Shashidhar, Dilek Hakkani-Tur
Assessment of reading comprehension through content-based interactions plays an important role in the reading acquisition process. In this paper, we propose a novel approach for ge…
Spark: A System for Scientifically Creative Idea Generation
Aishik Sanyal, Samuel Schapiro, Sumuk Shashidhar +3
Recently, large language models (LLMs) have shown promising abilities to generate novel research ideas in science, a direction which coincides with many foundational principles in…
YourBench: Easy Custom Evaluation Sets for Everyone
Sumuk Shashidhar, Clémentine Fourrier, Alina Lozovskia +3
Evaluating large language models (LLMs) effectively remains a critical bottleneck, as traditional static benchmarks suffer from saturation and contamination, while human evaluation…