activity
20242026
collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Cross-Domain Hybrid OPD for Generalizable Search Agents

Hongzhan Chen, Xiaoyu Liu, Dengming Zhang +11

Recent advances in Reinforcement Learning (RL) have substantially improved the capabilities of autonomous search agents, enabling sophisticated planning, and iterative retrieval ov…

cs.CL2026

Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization

Shiping Gao, Hongzhan Chen, Xiaojun Quan +2

Process reward models (PRMs) provide fine-grained supervision for reasoning, but reliable PRMs often require step annotations or heavy verification pipelines, making them costly to…

cs.CL2025

Discriminative Policy Optimization for Token-Level Reward Models

Hongzhan Chen, Tao Yang, Shiping Gao +4

Process reward models (PRMs) provide more nuanced supervision compared to outcome reward models (ORMs) for optimizing policy models, positioning them as a promising approach to enh…

cs.CL2024

Knowledge Distillation of Black-Box Large Language Models

Hongzhan Chen, Ruijun Chen, Yuqi Yi +4

Given the exceptional performance of proprietary large language models (LLMs) like GPT-4, recent research has increasingly focused on boosting the capabilities of smaller models th…

cs.CL2024

SocialBench: Sociality Evaluation of Role-Playing Conversational Agents

Hongzhan Chen, Hehong Chen, Ming Yan +8

Large language models (LLMs) have advanced the development of various AI conversational agents, including role-playing conversational agents that mimic diverse characters and human…