6 citations · 6 across the 1 of their papers we have counts for
4 papers
MS MARCO Web Search: a Large-scale Information-rich Web Dataset with Millions of Real Click Labels
Qi Chen, Xiubo Geng, Corby Rosset +28
Recent breakthroughs in large models have highlighted the critical significance of data scale, labels and modals. In this paper, we introduce MS MARCO Web Search, the first large-s…
Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for LLM Web Agents
Corby Rosset, Ho-Lam Chung, Guanghui Qin +5
Existing question answering (QA) datasets are no longer challenging to most powerful Large Language Models (LLMs). Traditional QA benchmarks like TriviaQA, NaturalQuestions, ELI5 a…
Dodo: Dynamic Contextual Compression for Decoder-only LMs
Guanghui Qin, Corby Rosset, Ethan C. Chau +2
Transformer-based language models (LMs) are inefficient in long contexts. We propose Dodo, a solution for context compression. Instead of one vector per token in a standard transfo…
Automatic Pair Construction for Contrastive Post-training
Canwen Xu, Corby Rosset, Ethan C. Chau +6
Alignment serves as an important step to steer large language models (LLMs) towards human preferences. In this paper, we propose an automatic way to construct contrastive data for…