2 papers
cs.LG2026
ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning
Davit Melikidze, Marian Schneider, Jessica Lam +4
Reinforcement Learning from Human Feedback (RLHF) has become the standard for aligning Large Language Models (LLMs), yet its efficacy is bottlenecked by the high cost of acquiring…
cs.HC2024
MODOC: A Modular Interface for Flexible Interlinking of Text Retrieval and Text Generation Functions
Yingqiang Gao, Jhony Prada, Nianlong Gu +2
Large Language Models (LLMs) produce eloquent texts but often the content they generate needs to be verified. Traditional information retrieval systems can assist with this task, b…