activity
20242026
most citedEfficient Rationale-based Retrieval: On-policy Distillation from Generative Rerankers based on JEPA

1 citations · 1 across the 4 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

ToSCA: Leveraging Hierarchical Reinforcement Learning on Temporal and Strategic Abstractions of Conversational Agents

Xiaoyu Wang, Qingqing Gu, Yue Zhao +5

Humans naturally exhibit multiple forms of abstraction in reasoning and interaction, including temporal abstraction across decision timescales and strategic abstraction over commun…

cs.CL2025

Chain-of-Conceptual-Thought Elicits Daily Conversation in Large Language Models

Qingqing Gu, Dan Wang, Yue Zhao +5

Chain-of-Thought (CoT) is widely applied to enhance the LLM capability in math, coding and reasoning tasks. However, its performance is limited for open-domain tasks, when there ar…

cs.CL2025

Dream to Chat: Model-based Reinforcement Learning on Dialogues with User Belief Modeling

Yue Zhao, Xiaoyu Wang, Dan Wang +7

World models have been widely utilized in robotics, gaming, and auto-driving. However, their applications on natural language tasks are relatively limited. In this paper, we constr…

cs.CL2025

Convert Language Model into a Value-based Strategic Planner

Xiaoyu Wang, Yue Zhao, Qingqing Gu +4

Emotional support conversation (ESC) aims to alleviate the emotional distress of individuals through effective conversations. Although large language models (LLMs) have obtained re…

cs.CL2025

EmoFSM: A Finite State Machine for Emotional Support Conversation

Yue Zhao, Qingqing Gu, Xiaoyu Wang +5

Emotional support conversation (ESC) aims to alleviate people's emotional distress through effective conversations. Although large language models (LLMs) have made remarkable progr…

cs.CL2024

Multi-Party Supervised Fine-tuning of Language Models for Multi-Party Dialogue Generation

Xiaoyu Wang, Ningyuan Xi, Teng Chen +6

Large Language Models (LLM) are usually fine-tuned to participate in dyadic or two-party dialogues, which can not adapt well to multi-party dialogues (MPD), which hinders their app…