6 papers
A Formula-Driven Survey and Research Agenda for On-Policy Distillation
Bowen Zhang
On-policy distillation (OPD) trains an LLM on states induced by the current or recent student policy: the student generates complete or partial rollouts, a teacher or self-teacher…
OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations
Jiangwang Chen, Bowen Zhang, Zixin Song +4
Although large language model (LLM) conversational systems process millions of multi-turn dialogues daily, they remain fundamentally reactive: they respond only after the user type…
Co-ReAct: Rubrics as Step-Level Collaborators for ReAct Agents
Jiazheng Kang, Bowen Zhang, Zixin Song +4
ReAct-style agents for search-intensive, multi-step reasoning tasks rely largely on their own internal judgment to decide what evidence to seek, which reasoning or action step to t…
PoLi-RL: A Point-to-List Reinforcement Learning Framework for Conditional Semantic Textual Similarity
Zixin Song, Bowen Zhang, Qian-Wen Zhang +3
Conditional Semantic Textual Similarity (C-STS) measures the semantic proximity between text segments under a specific condition, thereby overcoming the ambiguity inherent in tradi…
CoDiEmb: A Collaborative yet Distinct Framework for Unified Representation Learning in Information Retrieval and Semantic Textual Similarity
Bowen Zhang, Zixin Song, Chunquan Chen +3
Learning unified text embeddings that excel across diverse downstream tasks is a central goal in representation learning, yet negative transfer remains a persistent obstacle. This…
CSE-SFP: Enabling Unsupervised Sentence Representation Learning via a Single Forward Pass
Bowen Zhang, Zixin Song, Chunping Li
As a fundamental task in Information Retrieval and Computational Linguistics, sentence representation has profound implications for a wide range of practical applications such as t…