activity
20242026
most citedReinforcement Learning for LLM Post-Training: A Survey

9 citations · 10 across the 6 of their papers we have counts for

collaborators

11 papers

cs.MM2026

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models

Yuwen Wang, Tian-Hao Zhang, Minghao Cai +7

Complex acoustic problems may require models to perform acoustic operations, interact with external tools and reason over the resulting textual or processed-audio observations rath…

cs.RO2026

Labimus: A Simulation and Benchmark for Humanoid Dexterous Manipulation in Chemical Laboratory

Yuhan Wu, Zhao Jin, Tao Li +9

Laboratory automation has made remarkable progress through robotic platforms and AI-driven scientific reasoning. However, many laboratory operations (e.g., solid--solid transfer) r…

cs.CL20269 cited

Reinforcement Learning for LLM Post-Training: A Survey

Zhichao Wang, Kiran Ramnath, Bin Bi +9

Large language models (LLMs) trained via pretraining and supervised fine-tuning (SFT) can still produce harmful and misaligned outputs, or struggle in domains like math and coding.…

cs.LG2026

GIFT: Group-Relative Implicit Fine-Tuning Integrates GRPO with DPO and UNA

Zhichao Wang

This paper investigates whether reward matching is a viable alternative to reward maximization methods for on-policy RL of LLMs. Group-relative Implicit Fine-Tuning (GIFT) is propo…

cs.CL2026

UFT: Unifying Fine-Tuning of SFT and RLHF/DPO/UNA through a Generalized Implicit Reward Function

Zhichao Wang, Bin Bi, Zixu Zhu +6

By pretraining on trillions of tokens, an LLM gains the capability of text generation. However, to enhance its utility and reduce potential harm, SFT and alignment are applied sequ…

cs.LG20261 cited

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types

Zhichao Wang, Bin Bi, Can Huang +7

RL alignment methods, including RLHF and DPO, are primarily based on pairwise preference data. Although scalar or score-based feedback has been collected in some settings, it is ra…