collaborators

6 papers

cs.LG2025

ToolSample: Dual Dynamic Sampling Methods with Curriculum Learning for RL-based Tool Learning

Zihao Feng, Xiaoxue Wang, Bowen Wu +4

While reinforcement learning (RL) is increasingly used for LLM-based tool learning, its efficiency is often hampered by an overabundance of simple samples that provide diminishing…

cs.MA2025

Empowering LLMs in Task-Oriented Dialogues: A Domain-Independent Multi-Agent Framework and Fine-Tuning Strategy

Zihao Feng, Xiaoxue Wang, Bowen Wu +6

Task-oriented dialogue systems based on Large Language Models (LLMs) have gained increasing attention across various industries and achieved significant results. Current approaches…

cs.CL2025

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward

Zongsheng Wang, Kaili Sun, Bowen Wu +3

Role-playing conversational agents (RPCAs) face persistent challenges in maintaining role consistency. To address this, we propose RAIDEN-R1, a novel reinforcement learning framewo…

cs.SD2025

F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization

Xiaohui Sun, Ruitong Xiao, Jianye Mo +3

We present F5R-TTS, a novel text-to-speech (TTS) system that integrates Group Relative Policy Optimization (GRPO) into a flow-matching based architecture. By reformulating the dete…

cs.CL2025

Improving Generalization in Intent Detection: GRPO with Reward-Based Curriculum Sampling

Zihao Feng, Xiaoxue Wang, Ziwei Bai +4

Intent detection, a critical component in task-oriented dialogue (TOD) systems, faces significant challenges in adapting to the rapid influx of integrable tools with complex interr…

cs.CL2025

Interpersonal Memory Matters: A New Task for Proactive Dialogue Utilizing Conversational History

Bowen Wu, Wenqing Wang, Haoran Li +3

Proactive dialogue systems aim to empower chatbots with the capability of leading conversations towards specific targets, thereby enhancing user engagement and service autonomy. Ex…