collaborators

5 papers

cs.LG2026

Provably Efficient Policy-Reward Co-Pretraining for Adversarial Imitation Learning

Tian Xu, Zexuan Chen, Zhilong Zhang +4

Adversarial imitation learning (AIL) achieves high-quality imitation compared to behavioral cloning (BC), but demands substantial online environment interaction. Recent empirical w…

cs.LG2026

Non-Adversarial Imitation Learning Provably Free of Compounding Errors: The Value Flow Mechanism

Tian Xu, Chenyang Wang, Xiaochen Zhai +3

Adversarial imitation learning (AIL) achieves high-quality imitation by mitigating compounding errors inherent to behavioral cloning (BC), yet its adversarial optimization frequent…

cs.IR2026

K-CARE: Knowledge-driven Symmetrical Contextual Anchoring and Analogical Prototype Reasoning for E-commerce Relevance

Chen Yifei, Tian Zhixing, Wang Chenyang +1

This paper targets e-commerce search relevance. While Large Language Models (LLMs) have demonstrated significant potential in this field, they often encounter performance bottlenec…

cs.LG2026

Off-Policy Value-Based Reinforcement Learning for Large Language Models

Peng-Yuan Wang, Ziniu Li, Tian Xu +8

Improving data utilization efficiency is critical for scaling reinforcement learning (RL) for long-horizon tasks where generating trajectories is expensive. However, the dominant R…

cs.AI2025

A Survey on Large Language Models for Mathematical Reasoning

Peng-Yuan Wang, Tian-Shuo Liu, Chenyang Wang +8

Mathematical reasoning has long represented one of the most fundamental and challenging frontiers in artificial intelligence research. In recent years, large language models (LLMs)…