collaborators

7 papers

cs.AI2026

SkillEval: Decomposing Agent Skill Quality into Interpretable Signals

Jiahui Han, Qinuo Li, Ziheng Peng +6

Agent skills provide reusable procedural knowledge that helps agents solve specialized tasks. As their use expands, evaluating skill quality becomes increasingly important. Existin…

cs.CL2026

Self-Improving Large Language Models via Progressive Experience Evolution

Shijie Ren, Xiting Wang, Meng Li +8

Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism for transforming transient interaction expe…

cs.AI2026

Do LLMs Know Their Vulnerable Scenarios?

Ziheng Peng, Huiqi Deng, Haoran Jing +5

Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards. Existing red-teami…

cs.LG2026

To Call or Not to Call: Diagnosing Intrinsic Over-Calling Bias in LLM Agents

Wei Shi, Ziheng Peng, Sihang Li +4

LLM agents exhibit a consistent tendency to over-call, invoking tools even in situations where none is needed. On the When2Call benchmark, six models from three families show high…

cs.LG2026

Target Concept Tuning Improves Extreme Weather Forecasting

Shijie Ren, Xinyue Gu, Ziheng Peng +6

Deep learning models for meteorological forecasting often fail in rare but high-impact events such as typhoons, where relevant data is scarce. Existing fine-tuning methods typicall…

cs.LG2026

ProtoTS: Learning Hierarchical Prototypes for Explainable Time Series Forecasting

Ziheng Peng, Shijie Ren, Xinyue Gu +3

While deep learning has achieved impressive performance in time series forecasting, it becomes increasingly crucial to understand its decision-making process for building trust in…