most citedSecrets of RLHF in Large Language Models Part I: PPO

19 citations · 50 across the 10 of their papers we have counts for

collaborators

10 papers

cs.CL2024

In-Memory Learning: A Declarative Learning Framework for Large Language Models

Bo Wang, Tianxiang Sun, Hang Yan +3

The exploration of whether agents can align with their environment without relying on human-labeled data presents an intriguing research topic. Drawing inspiration from the alignme…

cs.CL20243 cited

Agent Alignment in Evolving Social Norms

Shimin Li, Tianxiang Sun, Qinyuan Cheng +1

Agents based on Large Language Models (LLMs) are increasingly permeating various domains of human production and life, highlighting the importance of aligning them with human value…

cs.LG20241 cited

Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT

Zhengfu He, Xuyang Ge, Qiong Tang +3

Sparse dictionary learning has been a rapidly growing technique in mechanistic interpretability to attack superposition and extract more human-understandable features from model ac…

cs.CL2024

LLM can Achieve Self-Regulation via Hyperparameter Aware Generation

Siyin Wang, Shimin Li, Tianxiang Sun +6

In the realm of Large Language Models (LLMs), users commonly employ diverse decoding strategies and adjust hyperparameters to control the generated text. However, a critical questi…

cs.CL2024

DenoSent: A Denoising Objective for Self-Supervised Sentence Representation Learning

Xinghao Wang, Junliang He, Pengyu Wang +3

Contrastive-learning-based methods have dominated sentence representation learning. These methods regularize the representation space by pulling similar sentence representations cl…

cs.CL20238 cited

Evaluating Hallucinations in Chinese Large Language Models

Qinyuan Cheng, Tianxiang Sun, Wenwei Zhang +8

In this paper, we establish a benchmark named HalluQA (Chinese Hallucination Question-Answering) to measure the hallucination phenomenon in Chinese large language models. HalluQA c…