collaborators

6 papers

cs.AI2026

Latent Thought Credit: Multi-Answer Credit Assignment for Latent Reasoning

Xuyang Zhao, Liting Zhang, Zichen Xu +4

Latent reasoning allows language models to carry out intermediate reasoning in continuous latent representations rather than fully externalizing it as discrete chains of thought. H…

cs.AI2026

Is More Privileged Information Better? From Solution Traces to Problem-Solving Structure in Self-Distilled Reasoning

Xuyang Zhao, Liting Zhang, Zichen Xu +4

On-policy self-distillation (OPSD) improves reasoning by using a privileged view of a model conditioned on reference solutions to supervise a student view that observes only the qu…

cs.CL2026

TARPO: Token-Wise Latent-Explicit Reasoning via Action-Routing Policy Optimization

Liting Zhang, Shiwan Zhao, Xuyang Zhao +3

Latent reasoning has emerged as a promising alternative to discrete Chain-of-Thought (CoT) in large language models (LLMs), enabling more expressive reasoning by operating over con…

cs.LG2026

Density-Guided Robust Counterfactual Explanations on Tabular Data under Model Multiplicity

Jun Tan, Qing Guo, Zicheng Xu +3

Counterfactual explanations (CEs) are essential for actionable recourse, yet their reliability is often compromised in low-density regions, where classifiers exhibit high variance.…

cs.CL2026

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning

Shiwan Zhao, Zhihu Wang, Xuyang Zhao +10

Post-training has become central to turning pretrained large language models (LLMs) into aligned, capable, and deployable systems. Recent progress spans supervised fine-tuning (SFT…

cs.CL2026

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

Feng Luo, Yu-Neng Chuang, Guanchu Wang +4

On-policy distillation (OPD) trains student models under their own induced distribution while leveraging supervision from stronger teachers. We identify a failure mode of OPD: as t…