5 papers
Experiment-as-Code Labs: A Declarative Stack for AI-Driven Scientific Discovery
Zhenning Yang, Yuhan Chen, Patrick Tser Jern Kon +5
To unleash the full potential of AI for Science, we must untether the agents from a purely digital environment. The agent's ability to control and explore in real-world labs is ess…
Contamination Means Overestimation? A Fine-Grained Empirical Study in Code Intelligence
Zhen Yang, Hongyi Lin, Yifan He +7
In recent years, code intelligence has gained increasing importance in the field of automated software engineering. Meanwhile, the widespread adoption of Pretrained Language Models…
R2ComSync: Improving Code-Comment Synchronization with In-Context Learning and Reranking
Zhen Yang, Hongyi Lin, Xiao Yu +5
Code-Comment Synchronization (CCS) aims to synchronize the comments with code changes in an automated fashion, thereby significantly reducing the workload of developers during soft…
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models
Xiaoyang Chen, Xinan Dai, Yu Du +28
To advance the mathematical proficiency of large language models (LLMs), the DeepMath team has launched an open-source initiative aimed at developing an open mathematical LLM and s…
Exploring and Lifting the Robustness of LLM-powered Automated Program Repair with Metamorphic Testing
Pengyu Xue, Linhao Wu, Zhen Yang +8
In recent years, Large language model-powered Automated Program Repair (LAPR) techniques have achieved state-of-the-art bug-fixing performance and have been pervasively applied and…