9 papers
PolyAlign: Conditional Human-Distribution Alignment
L. D. M. S. Sai Teja, Ufaq Khan, Sathira Silva +2
Post-training methods such as supervised fine-tuning (SFT) and preference optimization typically align language models toward a single global assistant behavior. While effective fo…
Beyond Output Critique: Self-Correction via Task Distillation
Hossein A. Rahmani, Mengting Wan, Pei Zhou +4
Large language models (LLMs) have shown promising self-correction abilities, where iterative refinement improves the quality of generated responses. However, most existing approach…
Fine-tuning Small Language Models as Efficient Enterprise Search Relevance Labelers
Yue Kang, Zhuoyi Huang, Benji Schussheim +19
In enterprise search, building high-quality datasets at scale remains a central challenge due to the difficulty of acquiring labeled data. To resolve this challenge, we propose an…
Overview of the TREC 2022 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz +4
This is the fourth year of the TREC Deep Learning track. As in previous years, we leverage the MS MARCO datasets that made hundreds of thousands of human annotated training labels…
Overview of the TREC 2023 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz +5
This is the fifth year of the TREC Deep Learning track. As in previous years, we leverage the MS MARCO datasets that made hundreds of thousands of human-annotated training labels a…
Overview of the TREC 2021 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz +2
This is the third year of the TREC Deep Learning track. As in previous years, we leverage the MS MARCO datasets that made hundreds of thousands of human annotated training labels a…