5 papers · 1 filter
Cross-Domain Hybrid OPD for Generalizable Search Agents
Hongzhan Chen, Xiaoyu Liu, Dengming Zhang +11
Recent advances in Reinforcement Learning (RL) have substantially improved the capabilities of autonomous search agents, enabling sophisticated planning, and iterative retrieval ov…
Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization
Shiping Gao, Hongzhan Chen, Xiaojun Quan +2
Process reward models (PRMs) provide fine-grained supervision for reasoning, but reliable PRMs often require step annotations or heavy verification pipelines, making them costly to…
Discriminative Policy Optimization for Token-Level Reward Models
Hongzhan Chen, Tao Yang, Shiping Gao +4
Process reward models (PRMs) provide more nuanced supervision compared to outcome reward models (ORMs) for optimizing policy models, positioning them as a promising approach to enh…
Knowledge Distillation of Black-Box Large Language Models
Hongzhan Chen, Ruijun Chen, Yuqi Yi +4
Given the exceptional performance of proprietary large language models (LLMs) like GPT-4, recent research has increasingly focused on boosting the capabilities of smaller models th…
SocialBench: Sociality Evaluation of Role-Playing Conversational Agents
Hongzhan Chen, Hehong Chen, Ming Yan +8
Large language models (LLMs) have advanced the development of various AI conversational agents, including role-playing conversational agents that mimic diverse characters and human…