most citedSecrets of RLHF in Large Language Models Part I: PPO

19 citations · 43 across the 13 of their papers we have counts for

collaborators
Showing cs.CLShow all

16 papers · 1 filter

cs.CL202319 cited

Secrets of RLHF in Large Language Models Part I: PPO

Rui Zheng, Shihan Dou, Songyang Gao +24

Large language models (LLMs) have formulated a blueprint for the advancement of artificial general intelligence. Its primary objective is to function as a human-centric (helpful, h…

cs.CL2023

Actively Supervised Clustering for Open Relation Extraction

Jun Zhao, Yongxin Zhang, Qi Zhang +4

Current clustering-based Open Relation Extraction (OpenRE) methods usually adopt a two-stage pipeline. The first stage simultaneously learns relation representations and assignment…

cs.CL20231 cited

RE-Matching: A Fine-Grained Semantic Matching Method for Zero-Shot Relation Extraction

Jun Zhao, Wenyu Zhan, Xin Zhao +6

Semantic matching is a mainstream paradigm of zero-shot relation extraction, which matches a given input with a corresponding label description. The entities in the input should ex…

cs.CL2023

Open Set Relation Extraction via Unknown-Aware Training

Jun Zhao, Xin Zhao, Wenyu Zhan +6

The existing supervised relation extraction methods have achieved impressive performance in a closed-set setting, where the relations during both training and testing remain the sa…

cs.CL2023

Farewell to Aimless Large-scale Pretraining: Influential Subset Selection for Language Model

Xiao Wang, Weikang Zhou, Qi Zhang +7

Pretrained language models have achieved remarkable success in various natural language processing tasks. However, pretraining has recently shifted toward larger models and larger…

cs.CL2023

Modeling the Q-Diversity in a Min-max Play Game for Robust Optimization

Ting Wu, Rui Zheng, Tao Gui +2

Models trained with empirical risk minimization (ERM) are revealed to easily rely on spurious correlations, resulting in poor generalization. Group distributionally robust optimiza…