2 papers
cs.CL2026
DUAL-Bench: Measuring Over-Refusal and Robustness in Vision-Language Models
Kaixuan Ren, Preslav Nakov, Usman Naseem
As vision-language models (VLMs) become increasingly capable, maintaining a balance between safety and usefulness remains a central challenge. Safety mechanisms, while essential, c…
cs.CL2025
TurnBench-MS: A Benchmark for Evaluating Multi-Turn, Multi-Step Reasoning in Large Language Models
Yiran Zhang, Mo Wang, Xiaoyang Li +3
Despite impressive advances in large language models (LLMs), existing benchmarks often focus on single-turn or single-step tasks, failing to capture the kind of iterative reasoning…