2 papers
cs.AI2025
ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
Tianyu Hua, Harper Hua, Violet Xiang +5
Large language models (LLMs) have shown promise in transforming machine learning research, yet their capability to faithfully implement novel ideas from recent research papers-idea…
cs.HC2024
ChatCollab: Exploring Collaboration Between Humans and AI Agents in Software Teams
Benjamin Klieger, Charis Charitsis, Miroslav Suzara +3
We explore the potential for productive team-based collaboration between humans and Artificial Intelligence (AI) by presenting and conducting initial tests with a general framework…