collaborators

10 papers

cs.AI2026

KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

Yaxing Lyu, Shengjie Zhou, Binbin Toh +2

As LLMs increasingly act through tools, they must reconcile user instructions, parametric knowledge, and dynamic environmental observations before taking actions. We introduce KC-B…

cs.AI2026

Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions

Pengyu Zhu, Lijun Li, Longju Yang +2

Deep Research agents conduct long-horizon investigations by iteratively planning, retrieving evidence, and generating reports. However, it remains unclear whether they can resist a…

cs.AI2026

SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment

Hao Li, Jingkun An, Zijun Song +8

Aligning Large Language Models (LLMs) with human values often degrades their general capabilities, termed the alignment tax. Existing methods mitigate this by balancing dual object…

cs.CL2026

STT-Arena: A More Realistic Environment for Tool-Using with Spatio-Temporal Dynamics

Tingfeng Hui, Hao Xu, Pengyu Zhu +5

Large language models (LLMs) deployed in real-world agentic applications must be capable of replanning and adapting when mid-task disruptions invalidate their prior decisions. Exis…

cs.AI2026

UniACE: A Unified Framework for Evaluating LLM Agentic Capabilities

Pengyu Zhu, Lijun Li, Yaxing Lyu +8

Agent benchmarks are increasingly used to compare large language models (LLMs) across domains, yet a reported score reflects a complete model--harness--environment configuration ra…

cs.AI2026

"LLM Agent Performance" Is Not a Single Evaluation Target

Pengyu Zhu, Li Sun, Philip S. Yu +1

LLM agent benchmark scores are shaped not only by the model but also by the agent harness, environment, evaluator, and inference budget. Unified execution controls these non-model…