collaborators

6 papers

cs.AI2026

Latent Thought Credit: Multi-Answer Credit Assignment for Latent Reasoning

Xuyang Zhao, Liting Zhang, Zichen Xu +4

Latent reasoning allows language models to carry out intermediate reasoning in continuous latent representations rather than fully externalizing it as discrete chains of thought. H…

cs.AI2026

Is More Privileged Information Better? From Solution Traces to Problem-Solving Structure in Self-Distilled Reasoning

Xuyang Zhao, Liting Zhang, Zichen Xu +4

On-policy self-distillation (OPSD) improves reasoning by using a privileged view of a model conditioned on reference solutions to supervise a student view that observes only the qu…

cs.CL2026

TARPO: Token-Wise Latent-Explicit Reasoning via Action-Routing Policy Optimization

Liting Zhang, Shiwan Zhao, Xuyang Zhao +3

Latent reasoning has emerged as a promising alternative to discrete Chain-of-Thought (CoT) in large language models (LLMs), enabling more expressive reasoning by operating over con…

cs.CL2026

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning

Shiwan Zhao, Zhihu Wang, Xuyang Zhao +10

Post-training has become central to turning pretrained large language models (LLMs) into aligned, capable, and deployable systems. Recent progress spans supervised fine-tuning (SFT…

cs.CL2025

MAPEX: A Multi-Agent Pipeline for Keyphrase Extraction

Liting Zhang, Shiwan Zhao, Aobo Kong +1

Keyphrase extraction is a fundamental task in natural language processing. However, existing unsupervised prompt-based methods for Large Language Models (LLMs) often rely on single…

cs.AI2025

AgentCDM: Enhancing Multi-Agent Collaborative Decision-Making via ACH-Inspired Structured Reasoning

Xuyang Zhao, Shiwan Zhao, Hualong Yu +2

Multi-agent systems (MAS) powered by large language models (LLMs) hold significant promise for solving complex decision-making tasks. However, the core process of collaborative dec…