collaborators

6 papers

cs.LG2026

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment

Yang Tian, Rui Wang, Xumeng Wen +5

Long-horizon agentic tasks pose a fundamental credit assignment challenge for outcome-base reinforcement learning: trajectory-level rewards verify final correctness but provide lim…

cs.CL2026

Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability

Yang Tian, Zhengpeng Shi, Yu Zhou +1

Large language models are increasingly deployed as agents that solve tasks by interacting with external tool environments. Although recent tool-use benchmarks increasingly cover co…

cs.RO2026

LA4VLA: Learning to Act without Seeing via Language-Action Pretraining

Tao Lin, Yuxin Du, Yiran Mao +13

Vision-Language-Action (VLA) models are commonly pretrained on robot demonstrations by jointly mapping visual observations and language instructions to actions. However, dense visu…

cs.AI2026

From Digital to Physical: Digital Agents as Autonomous Coaches for Physical Intelligence

Zixing Lei, Genjia Liu, Yuanshuo Zhang +11

The field of Embodied AI is witnessing a rapid evolution toward general-purpose robotic systems, fueled by high-fidelity simulation and large-scale data collection. However, this s…

cs.AI2026

MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs

Huining Yuan, Zelai Xu, Zheyue Tan +10

Developing Large Language Models (LLMs) to cooperate and compete effectively within multi-agent systems (MASs) is a critical step towards more advanced intelligence. While reinforc…

cs.DC2025

EARL: Efficient Agentic Reinforcement Learning Systems for Large Language Models

Zheyue Tan, Mustapha Abdullahi, Tuo Shi +5

Reinforcement learning (RL) has become a pivotal component of large language model (LLM) post-training, and agentic RL extends this paradigm to operate as agents through multi-turn…