2 citations · 2 across the 2 of their papers we have counts for
4 papers
Explainable reinforcement learning from human feedback to improve alignment
Shicheng Liu, Siyuan Xu, Wenjie Qiu +2
A common and effective strategy for humans to improve an unsatisfactory outcome in daily life is to find a cause of this outcome and correct the cause. In this paper, we investigat…
The Path of Self-Evolving Large Language Models: Achieving Data-Efficient Learning via Intrinsic Feedback
Hangfan Zhang, Siyuan Xu, Zhimeng Guo +8
Reinforcement learning (RL) has demonstrated potential in enhancing the reasoning capabilities of large language models (LLMs), but such training typically demands substantial effo…
Simple Denoising Diffusion Language Models
Huaisheng Zhu, Zhengyu Chen, Shijie Zhou +8
Recent Uniform State Diffusion Models (USDMs), initialized from a uniform prior, offer the promise of fast text generation due to their inherent self-correction ability compared to…
Beyond GPT-5: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing
Yiqun Zhang, Hao Li, Jianhao Chen +4
Balancing performance and efficiency is a central challenge in large language model (LLM) advancement. GPT-5 addresses this with test-time routing, dynamically assigning queries to…