collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2025

MRO: Enhancing Reasoning in Diffusion Language Models via Multi-Reward Optimization

Chenglong Wang, Yang Gan, Hang Zhou +10

Recent advances in diffusion language models (DLMs) have presented a promising alternative to traditional autoregressive large language models (LLMs). However, DLMs still lag behin…

cs.CL2025

GRAM-R: Self-Training Generative Foundation Reward Models for Reward Reasoning

Chenglong Wang, Yongyu Mu, Hang Zhou +10

Significant progress in reward modeling over recent years has been driven by a paradigm shift from task-specific designs towards generalist reward models. Despite this trend, devel…

cs.CL2025

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models

Kaiyan Chang, Yonghao Shi, Chenglong Wang +7

Test-Time Scaling (TTS) is a promising approach to progressively elicit the model's intelligence during inference. Recently, training-based TTS methods, such as continued reinforce…

cs.CL20241 cited

Hybrid Alignment Training for Large Language Models

Chenglong Wang, Hang Zhou, Kaiyan Chang +5

Alignment training is crucial for enabling large language models (LLMs) to cater to human intentions and preferences. It is typically performed based on two stages with different o…

cs.CL2024

Prior Constraints-based Reward Model Training for Aligning Large Language Models

Hang Zhou, Chenglong Wang, Yimin Hu +3

Reinforcement learning with human feedback for aligning large language models (LLMs) trains a reward model typically using ranking loss with comparison pairs.However, the training…

cs.CL2023

Learning Evaluation Models from Large Language Models for Sequence Generation

Chenglong Wang, Hang Zhou, Kaiyan Chang +6

Automatic evaluation of sequence generation, traditionally reliant on metrics like BLEU and ROUGE, often fails to capture the semantic accuracy of generated text sequences due to t…