10 papers
KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents
Yaxing Lyu, Shengjie Zhou, Binbin Toh +2
As LLMs increasingly act through tools, they must reconcile user instructions, parametric knowledge, and dynamic environmental observations before taking actions. We introduce KC-B…
Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions
Pengyu Zhu, Lijun Li, Longju Yang +2
Deep Research agents conduct long-horizon investigations by iteratively planning, retrieving evidence, and generating reports. However, it remains unclear whether they can resist a…
SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
Hao Li, Jingkun An, Zijun Song +8
Aligning Large Language Models (LLMs) with human values often degrades their general capabilities, termed the alignment tax. Existing methods mitigate this by balancing dual object…
STT-Arena: A More Realistic Environment for Tool-Using with Spatio-Temporal Dynamics
Tingfeng Hui, Hao Xu, Pengyu Zhu +5
Large language models (LLMs) deployed in real-world agentic applications must be capable of replanning and adapting when mid-task disruptions invalidate their prior decisions. Exis…
UniACE: A Unified Framework for Evaluating LLM Agentic Capabilities
Pengyu Zhu, Lijun Li, Yaxing Lyu +8
Agent benchmarks are increasingly used to compare large language models (LLMs) across domains, yet a reported score reflects a complete model--harness--environment configuration ra…
"LLM Agent Performance" Is Not a Single Evaluation Target
Pengyu Zhu, Li Sun, Philip S. Yu +1
LLM agent benchmark scores are shaped not only by the model but also by the agent harness, environment, evaluator, and inference budget. Unified execution controls these non-model…