4 citations · 8 across the 4 of their papers we have counts for
6 papers · 1 filter
The Trickle-down Impact of Reward (In-)consistency on RLHF
Lingfeng Shen, Sihao Chen, Linfeng Song +5
Standard practice within Reinforcement Learning from Human Feedback (RLHF) involves optimizing against a Reward Model (RM), which itself is trained to reflect human preferences for…
Stabilizing RLHF through Advantage Model and Selective Rehearsal
Baolin Peng, Linfeng Song, Ye Tian +3
Large Language Models (LLMs) have revolutionized natural language processing, yet aligning these models with human values and preferences using RLHF remains a significant challenge…
Salience Allocation as Guidance for Abstractive Summarization
Fei Wang, Kaiqiang Song, Hongming Zhang +6
Abstractive summarization models typically learn to capture the salient information from scratch implicitly. Recent literature adds extractive summaries as guidance for abstractive…
The Importance of Category Labels in Grammar Induction with Child-directed Utterances
Lifeng Jin, William Schuler
Recent progress in grammar induction has shown that grammar induction is possible without explicit assumptions of language-specific knowledge. However, evaluation of induced gramma…
Depth-bounding is effective: Improvements and evaluation of unsupervised PCFG induction
Lifeng Jin, Finale Doshi-Velez, Timothy Miller +2
There have been several recent attempts to improve the accuracy of grammar induction systems by bounding the recursive complexity of the induction model (Ponvert et al., 2011; Noji…
Unsupervised Grammar Induction with Depth-bounded PCFG
Lifeng Jin, Finale Doshi-Velez, Timothy Miller +2
There has been recent interest in applying cognitively or empirically motivated bounds on recursion depth to limit the search space of grammar induction models (Ponvert et al., 201…