activity
20222025
most citedEditEval: An Instruction-Based Benchmark for Text Improvements

8 citations · 11 across the 2 of their papers we have counts for

collaborators
Showing cs.CLShow all

11 papers · 1 filter

cs.CL2025

Separate the Wheat from the Chaff: Winnowing Down Divergent Views in Retrieval Augmented Generation

Song Wang, Zihan Chen, Peng Wang +5

Retrieval-augmented generation (RAG) enhances large language models (LLMs) by integrating external knowledge sources to address their limitations in accessing up-to-date or special…

cs.CL2024

Source2Synth: Synthetic Data Generation and Curation Grounded in Real Data Sources

Alisia Lupidi, Carlos Gemmell, Nicola Cancedda +5

Synthetic data generation has recently emerged as a promising approach for enhancing the capabilities of large language models (LLMs) without the need for expensive human annotatio…

cs.CL20243 cited

Self-Taught Evaluators

Tianlu Wang, Ilia Kulikov, Olga Golovneva +7

Model-based evaluation is at the heart of successful model development -- as a reward model for training, and as a replacement for human evaluation. To train such evaluators, the s…

cs.CL2024

FairPair: A Robust Evaluation of Biases in Language Models through Paired Perturbations

Jane Dwivedi-Yu, Raaz Dwivedi, Timo Schick

The accurate evaluation of differential treatment in language models to specific groups is critical to ensuring a positive and safe user experience. An ideal evaluation should have…

cs.CL2024

TOOLVERIFIER: Generalization to New Tools via Self-Verification

Dheeraj Mekala, Jason Weston, Jack Lanchantin +4

Teaching language models to use tools is an important milestone towards building general assistants, but remains an open problem. While there has been significant progress on learn…

cs.CL2024

MultiContrievers: Analysis of Dense Retrieval Representations

Seraphina Goldfarb-Tarrant, Pedro Rodriguez, Jane Dwivedi-Yu +1

Dense retrievers compress source documents into (possibly lossy) vector representations, yet there is little analysis of what information is lost versus preserved, and how it affec…