2 papers
cs.SE2025
A Multi-Language Object-Oriented Programming Benchmark for Large Language Models
Shuai Wang, Liang Ding, Li Shen +4
Establishing fair and robust benchmarks is essential for evaluating intelligent code generation by large language models (LLMs). Our survey of 35 existing benchmarks uncovers three…
cs.CV2025
SRD: Reinforcement-Learned Semantic Perturbation for Backdoor Defense in VLMs
Shuhan Xu, Siyuan Liang, Hongling Zheng +6
Visual language models (VLMs) have made significant progress in image captioning tasks, yet recent studies have found they are vulnerable to backdoor attacks. Attackers can inject…