collaborators

5 papers

cs.LG2026

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment

Yang Tian, Rui Wang, Xumeng Wen +5

Long-horizon agentic tasks pose a fundamental credit assignment challenge for outcome-base reinforcement learning: trajectory-level rewards verify final correctness but provide lim…

cs.LG2025

Are Time-Series Foundation Models Deployment-Ready? A Systematic Study of Adversarial Robustness Across Domains

Jiawen Zhang, Zhenwei Zhang, Shun Zheng +3

Time-Series Foundation Models (TSFMs) are rapidly transitioning from research prototypes to core components of critical decision-making systems, driven by their impressive zero-sho…

cs.CL2025

Deep Self-Evolving Reasoning

Zihan Liu, Shun Zheng, Xumeng Wen +3

Long-form chain-of-thought reasoning has become a cornerstone of advanced reasoning in large language models. While recent verification-refinement frameworks have enabled proprieta…

cs.AI2025

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Xumeng Wen, Zihan Liu, Shun Zheng +9

Recent advancements in long chain-of-thought (CoT) reasoning, particularly through the Group Relative Policy Optimization algorithm used by DeepSeek-R1, have led to significant int…

cs.CL2025

Scalable In-Context Learning on Tabular Data via Retrieval-Augmented Large Language Models

Xumeng Wen, Shun Zheng, Zhen Xu +2

Recent studies have shown that large language models (LLMs), when customized with post-training on tabular data, can acquire general tabular in-context learning (TabICL) capabiliti…