3 papers
cs.SE2025
CoCo-Bench: A Comprehensive Code Benchmark For Multi-task Large Language Model Evaluation
Wenjing Yin, Tianze Sun, Yijiong Yu +19
Large language models (LLMs) play a crucial role in software engineering, excelling in tasks like code generation and maintenance. However, existing benchmarks are often narrow in…
cs.CL2024
Efficiently Quantifying and Mitigating Ripple Effects in Model Editing
Jianchen Wang, Zhouhong Gu, Xiaoxuan Zhu +5
Large Language Models have revolutionized numerous tasks with their remarkable efficacy. However, editing these models, crucial for rectifying outdated or erroneous information, of…
cs.CL2023
Piecing Together Clues: A Benchmark for Evaluating the Detective Skills of Large Language Models
Zhouhong Gu, Lin Zhang, Jiangjie Chen +8
Detectives frequently engage in information detection and reasoning simultaneously when making decisions across various cases, especially when confronted with a vast amount of info…