2 papers
cs.SE2026
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility
Jialun Cao, Yuk-Kit Chan, Zixuan Ling +12
Code-related benchmarks play a critical role in evaluating large language models (LLMs), yet their quality fundamentally shapes how the community interprets model capabilities. In…
cs.SE2025
AL-Bench: A Benchmark for Automatic Logging
Boyin Tan, Junjielong Xu, Zhouruixing Zhu +1
Logging, the practice of inserting log statements into source code, is critical for improving software reliability. Recently, language model-based techniques have been developed to…