◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Guangju Wang

3 papers hereh-index 3400 citations3 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author3

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.CL1
  • cs.DC1
  • cs.LG1

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.DC2025

ReaL: Efficient RLHF Training of Large Language Models with Parameter Reallocation

Zhiyu Mei, Wei Fu, Kaiwei Li +3

Reinforcement Learning from Human Feedback (RLHF) is a pivotal technique for empowering large language model (LLM) applications. Compared with the supervised training process of LL…

cs.LG2024

On Designing Effective RL Reward at Training Time for LLM Reasoning

Jiaxuan Gao, Shusheng Xu, Wenjie Ye +6

Reward models have been increasingly critical for improving the reasoning capability of LLMs. Existing research has shown that a well-trained reward model can substantially improve…

cs.CL2024

Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Shusheng Xu, Wei Fu, Jiaxuan Gao +6

Reinforcement Learning from Human Feedback (RLHF) is currently the most widely used method to align large language models (LLMs) with human preferences. Existing RLHF methods can b…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.