◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Zixu Zhu

3 papers hereh-index 3144 citations3 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author1

Across the 1 of 3 papers where every author was matched, so the position is known.

fields
  • cs.CL2
  • cs.LG1

identity via Semantic Scholar / OpenAlex

most citedReinforcement Learning for LLM Post-Training: A Survey

9 citations · 10 across the 3 of their papers we have counts for

collaborators

3 papers

cs.CL2024

UFT: Unifying Fine-Tuning of SFT and RLHF/DPO/UNA through a Generalized Implicit Reward Function

Zhichao Wang, Bin Bi, Zixu Zhu +6

By pretraining on trillions of tokens, an LLM gains the capability of text generation. However, to enhance its utility and reduce potential harm, SFT and alignment are applied sequ…

cs.LG2024★ 1 cited

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types

Zhichao Wang, Bin Bi, Can Huang +7

RL alignment methods, including RLHF and DPO, are primarily based on pairwise preference data. Although scalar or score-based feedback has been collected in some settings, it is ra…

cs.CL2024★ 9 cited

Reinforcement Learning for LLM Post-Training: A Survey

Zhichao Wang, Kiran Ramnath, Bin Bi +9

Large language models (LLMs) trained via pretraining and supervised fine-tuning (SFT) can still produce harmful and misaligned outputs, or struggle in domains like math and coding.…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.