collaborators

11 papers

cs.CL2026

OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills

Qiyuan Liu, Tingfeng Hui, Kun Zhan +2

LLM-based agents leverage third-party skills to extend their capabilities in open-world scenarios. However, third-party skills can introduce extra security vulnerabilities, as seem…

cs.LG2026

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition

Ningyuan Xi, Hao Xu, Hongsheng Xin +1

Large language models (LLMs) have made remarkable progress in reasoning tasks, largely driven by post-training paradigms, especially reinforcement learning with verifiable rewards…

cs.CL2026

Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions

Xuan Yang, Hao Xu, Tingfeng Hui +4

Despite great advances in tool-use capabilities of large language models (LLMs), existing evaluation benchmarks struggle to fully align with real-world scenarios. Such benchmarks m…

cs.LG2026

Linear Dynamics in the RLVR Training of Large Language Models

Tianle Wang, Jiayu Liu, Zhongyuan Wu +4

Reinforcement learning with verifiable rewards (RLVR) has driven significant performance gains in reasoning-oriented large language models (LLMs), yet its internal training dynamic…

cs.CL2026

STT-Arena: A More Realistic Environment for Tool-Using with Spatio-Temporal Dynamics

Tingfeng Hui, Hao Xu, Pengyu Zhu +5

Large language models (LLMs) deployed in real-world agentic applications must be capable of replanning and adapting when mid-task disruptions invalidate their prior decisions. Exis…

cs.LG2026

Verifier-Backed Hard Problem Generation for Mathematical Reasoning

Yuhang Lai, Jiazhan Feng, Yee Whye Teh +1

Large Language Models (LLMs) demonstrate strong capabilities for solving scientific and mathematical problems, yet they struggle to produce valid, challenging, and novel problems -…