3 papers
cs.LG2025
ResearchGPT: Benchmarking and Training LLMs for End-to-End Computer Science Research Workflows
Penghao Wang, Yuhao Zhou, Mengxuan Wu +12
As large language models (LLMs) advance, the ultimate vision for their role in science is emerging: we could build an AI collaborator to effectively assist human beings throughout…
cs.CL2025
Efficient Code LLM Training via Distribution-Consistent and Diversity-Aware Data Selection
Weijie Lyu, Sheng-Jun Huang, Xuan Xia
Recent advancements in large language models (LLMs) have significantly improved code generation and program comprehension, accelerating the evolution of software engineering. Curre…
cs.CL2025
Data-efficient LLM Fine-tuning for Code Generation
Weijie Lv, Xuan Xia, Sheng-Jun Huang
Large language models (LLMs) have demonstrated significant potential in code generation tasks. However, there remains a performance gap between open-source and closed-source models…