9 papers
A Unified Framework for the Evaluation of LLM Agentic Capabilities
Pengyu Zhu, Lijun Li, Yaxing Lyu +8
As LLMs are increasingly deployed as agents, reliable assessment of their agentic capabilities has become essential. However, reported benchmark scores often jointly reflect model…
From "Aha Moments" to Controllable Thinking: Toward Meta-Cognitive Reasoning in Large Reasoning Models via Decoupled Reasoning and Control
Rui Ha, Rui Pu, Chaozhuo Li +2
Large Reasoning Models (LRMs) can exhibit step-by-step reasoning, reflection, and backtracking, but these behaviors are often unregulated, leading to overthinking. As a result, LRM…
Structure-Guided Visual Perturbation Neutralization for LVLMs
Yuanhe Zhang, Xueting Wang, YanBin Ren +6
Image inputs enable Large Vision Language Models (LVLMs) to perceive fine-grained visual information, but also introduce a pixel-level attack surface through which adversarial pert…
"LLM Agent Performance" Is Not a Single Evaluation Target
Pengyu Zhu, Li Sun, Philip S. Yu +1
LLM agent benchmark scores are shaped not only by the model but also by the agent harness, environment, evaluator, and inference budget. Unified execution controls these non-model…
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation
Yi Liu, TingFeng Hui, Wei Zhang +4
Scalable AI agents training relies on interactive environments that faithfully simulate the consequences of agent actions. Manually crafted environments are expensive to build, bri…
Resource Consumption Threats in Large Language Models
Yuanhe Zhang, Xinyue Wang, Zhican Chen +8
Given limited and costly computational infrastructure, resource efficiency is a key requirement for large language models (LLMs). Efficient LLMs increase service capacity for provi…