6 papers
A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation
Hosna Oyarhoseini, Jimmy Lin, Amir-Hossein Karimi
Evaluation leaderboards such as LMArena play a central role in benchmarking large language models by aggregating pairwise human preferences into model rankings, yet the robustness…
NanoKnow: How to Know What Your Language Model Knows
Lingwei Gu, Nour Jedidi, Jimmy Lin
How do large language models (LLMs) know what they know? Answering this question has been difficult because pre-training data is often a "black box" - unknown or inaccessible. The…
A Systematic Study of Pseudo-Relevance Feedback with LLMs
Nour Jedidi, Jimmy Lin
Pseudo-relevance feedback (PRF) methods built on large language models (LLMs) can be organized along two key design dimensions: the feedback source, which is where the feedback tex…
Revisiting Feedback Models for HyDE
Nour Jedidi, Jimmy Lin
Recent approaches that leverage large language models (LLMs) for pseudo-relevance feedback (PRF) have generally not utilized well-established feedback models like Rocchio and RM3 w…
Study on LLMs for Promptagator-Style Dense Retriever Training
Daniel Gwon, Nour Jedidi, Jimmy Lin
Promptagator demonstrated that Large Language Models (LLMs) with few-shot prompts can be used as task-specific query generators for fine-tuning domain-specialized dense retrieval m…
Don't "Overthink" Passage Reranking: Is Reasoning Truly Necessary?
Nour Jedidi, Yung-Sung Chuang, James Glass +1
With the growing success of reasoning models across complex natural language tasks, researchers in the Information Retrieval (IR) community have begun exploring how similar reasoni…