6 papers
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Tong Zheng, Hongming Zhang, Wenhao Yu +7
Parallel thinking has emerged as a novel approach for enhancing the reasoning capabilities of large language models (LLMs) by exploring multiple reasoning paths concurrently. Howev…
Proactive Guidance of Multi-Turn Conversation in Industrial Search
Xiaoyu Li, Xiao Li, Li Gao +5
The evolution of Large Language Models (LLMs) has significantly advanced multi-turn conversation systems, emphasizing the need for proactive guidance to enhance users' interactions…
Free Lunch for User Experience: Crowdsourcing Agents for Scalable User Studies
Siyang Liu, Sahand Sabour, Xiaoyang Wang +1
User studies are central to user experience research, yet recruiting participant is expensive, slow, and limited in diversity. Recent work has explored using Large Language Models…
Enter the Void - Planning to Seek Entropy When Reward is Scarce
Ashish Sundar, Chunbo Luo, Xiaoyang Wang
Model-based reinforcement learning (MBRL) offers an intuitive way to increase the sample efficiency of model-free RL methods by simultaneously training a world model that learns to…
OpenCharacter: Training Customizable Role-Playing LLMs with Large-Scale Synthetic Personas
Xiaoyang Wang, Hongming Zhang, Tao Ge +3
Customizable role-playing in large language models (LLMs), also known as character generalization, is gaining increasing attention for its versatility and cost-efficiency in develo…
Addressing the sustainable AI trilemma: a case study on LLM agents and RAG
Hui Wu, Xiaoyang Wang, Zhong Fan
Large language models (LLMs) have demonstrated significant capabilities, but their widespread deployment and more advanced applications raise critical sustainability challenges, pa…