3 papers
cs.AI2026
A Unified Framework for the Evaluation of LLM Agentic Capabilities
Pengyu Zhu, Lijun Li, Yaxing Lyu +8
As LLMs are increasingly deployed as agents, reliable assessment of their agentic capabilities has become essential. However, reported benchmark scores often jointly reflect model…
cs.AI2026
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation
Yi Liu, TingFeng Hui, Wei Zhang +4
Scalable AI agents training relies on interactive environments that faithfully simulate the consequences of agent actions. Manually crafted environments are expensive to build, bri…
cs.CL2024
Alignment-Enhanced Decoding:Defending via Token-Level Adaptive Refining of Probability Distributions
Quan Liu, Zhenhong Zhou, Longzhu He +3
Large language models are susceptible to jailbreak attacks, which can result in the generation of harmful content. While prior defenses mitigate these risks by perturbing or inspec…