Showing cs.SEShow all
2 papers · 1 filter
cs.SE2025
Large Language Models are Qualified Benchmark Builders: Rebuilding Pre-Training Datasets for Advancing Code Intelligence Tasks
Kang Yang, Xinjun Mao, Shangwen Wang +7
Pre-trained code models rely heavily on high-quality pre-training data, particularly human-written reference comments that bridge code and natural language. However, these comments…
cs.SE2024
Fault Localization from the Semantic Code Search Perspective
Yihao Qin, Shangwen Wang, Yan Lei +5
The software development process is characterized by an iterative cycle of continuous functionality implementation and debugging, essential for the enhancement of software quality…