6 papers
Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth
Gaurav Sahu, Laurent Charlin, Christopher Pal
We study large-scale literature search from two complementary angles: improving the retrieval pipeline, and stress-testing the human reference list as an evaluation target. First,…
AInstein: Can LLMs Solve Research Problems From Parametric Memory Alone?
Shambhavi Mishra, Gaurav Sahu, Marco Pedersoli +3
Can large language models solve AI research problems using only their parametric knowledge, without fine-tuning, retrieval, or other external aids? We introduce AInstein, a framewo…
ReviewerToo: Should AI Join The Program Committee? A Look At The Future of Peer Review
Gaurav Sahu, Hugo Larochelle, Laurent Charlin +1
Peer review is the cornerstone of scientific publishing, yet it suffers from inconsistencies, reviewer subjectivity, and scalability challenges. We introduce ReviewerToo, a modular…
LitLLMs, LLMs for Literature Review: Are we there yet?
Shubham Agarwal, Gaurav Sahu, Abhay Puri +5
Literature reviews are an essential component of scientific research, but they remain time-intensive and challenging to write, especially due to the recent influx of research paper…
InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation
Gaurav Sahu, Abhay Puri, Juan Rodriguez +11
Data analytics is essential for extracting valuable insights from data that can assist organizations in making effective decisions. We introduce InsightBench, a benchmark dataset w…
A Guide To Effectively Leveraging LLMs for Low-Resource Text Summarization: Data Augmentation and Semi-supervised Approaches
Gaurav Sahu, Olga Vechtomova, Issam H. Laradji
Existing approaches for low-resource text summarization primarily employ large language models (LLMs) like GPT-3 or GPT-4 at inference time to generate summaries directly; however,…